S0136RS3-T15-S0136-Z · Full risk code

Node Failure Triggering Cascading Collapse

節點故障引發連鎖崩潰

Operation & Monitoring
Risk Description

When during peak hours a single node exits due to insufficient memory, and its traffic shifts to neighboring nodes which then overload in succession; due to unmitigated control gaps, the cascade spreads across the cluster within minutes, causing complete service outage, and the system collapses again after restart due to backlogged traffic, triggering compliance exposure and operational reputational costs.

Framework Mappings

OWASP Top 10 for LLMLLM10
MITRE ATLASAML.T0029
ISO/IEC 5338運作與監控
MIT AI Risk RepositoryDomain 7

Risk Treatment & Implementation Guidance

Apply circuit breakers so overloaded nodes fail fast instead of cascading; Design graceful degradation—queuing and simplified responses—to sustain basic service at peak; Capacity-plan and stress-test against peak scenarios, validating cluster failure behavior