S0072RS3-T09-S0072-Z · Full risk code

Adversarial Text Evading Detection

文字對抗樣本規避偵測

Verification & Validation
Risk Description

When malicious content processed through character manipulation successfully bypasses content filtering and enters the platform in volume; due to unmitigated control gaps, on manual review the content appears entirely ordinary to readers, yet the filtering rules were never triggered, triggering compliance exposure and operational reputational costs.

Framework Mappings

MITRE ATLASAML.T0043
NIST AI 100-2Evasion
OWASP Top 10 for LLMLLM01
ISO/IEC 5338驗證與確效
MIT AI Risk RepositoryDomain 2

Risk Treatment & Implementation Guidance

Normalize text before filtering, restoring character deformations, homoglyphs and inserted symbols; Deploy character-level anomaly detection identifying deliberate evasion patterns; Continuously update filters with known evasion samples and run red-team tests