Node Failure Triggering Cascading Collapse
節點故障引發連鎖崩潰
Operation & Monitoring
Risk Description
When during peak hours a single node exits due to insufficient memory, and its traffic shifts to neighboring nodes which then overload in succession; due to unmitigated control gaps, the cascade spreads across the cluster within minutes, causing complete service outage, and the system collapses again after restart due to backlogged traffic, triggering compliance exposure and operational reputational costs.
Framework Mappings
OWASP Top 10 for LLMLLM10
MITRE ATLASAML.T0029
ISO/IEC 5338運作與監控
MIT AI Risk RepositoryDomain 7
Risk Treatment & Implementation Guidance
Apply circuit breakers so overloaded nodes fail fast instead of cascading; Design graceful degradation—queuing and simplified responses—to sustain basic service at peak; Capacity-plan and stress-test against peak scenarios, validating cluster failure behavior