Adversarial Text Evading Detection
文字對抗樣本規避偵測
Verification & Validation
Risk Description
When malicious content processed through character manipulation successfully bypasses content filtering and enters the platform in volume; due to unmitigated control gaps, on manual review the content appears entirely ordinary to readers, yet the filtering rules were never triggered, triggering compliance exposure and operational reputational costs.
Framework Mappings
MITRE ATLASAML.T0043
NIST AI 100-2Evasion
OWASP Top 10 for LLMLLM01
ISO/IEC 5338驗證與確效
MIT AI Risk RepositoryDomain 2
Risk Treatment & Implementation Guidance
Normalize text before filtering, restoring character deformations, homoglyphs and inserted symbols; Deploy character-level anomaly detection identifying deliberate evasion patterns; Continuously update filters with known evasion samples and run red-team tests