S0121RS3-T10-S0121-Z · Full risk code

Multi-Step Inducement Attack Chain

多步驟誘導攻擊鏈

Verification & Validation
Risk Description

When an attacker advances through dozens of incremental conversational turns, each escalating only slightly from the last; due to unmitigated control gaps, judged turn by turn, the system finds every request reasonable, and ultimately produces content that was explicitly prohibited, triggering compliance exposure and operational reputational costs.

Framework Mappings

OWASP Top 10 for LLMLLM01
MITRE ATLASAML.T0051
NIST AI 600-1Information Security
ISO/IEC 5338運作與監控
MIT AI Risk RepositoryDomain 2

Risk Treatment & Implementation Guidance

Monitor multi-turn safety, assessing the trajectory of whole conversations rather than single turns; Score cumulative risk so progressively escalating request chains are refused and reset at threshold; Continuously drill with known attack chains to update cross-turn detection