S0070RS3-T09-S0070-Z · Full risk code

Adversarial Perturbation of Images

影像對抗擾動攻擊

Verification & Validation
Risk Description

When an attacker adds a crafted perturbation to input images, causing a detection system to classify an item that should be intercepted as normal; due to unmitigated control gaps, the system output gives no indication of anomaly, and human review cannot detect the perturbation because it is imperceptible, triggering compliance exposure and operational reputational costs.

Framework Mappings

MITRE ATLASAML.T0043
NIST AI 100-2Evasion
OWASP Top 10 for LLMLLM01
ISO/IEC 5338驗證與確效
MIT AI Risk RepositoryDomain 2

Risk Treatment & Implementation Guidance

Harden the model against perturbed inputs with adversarial training; Deploy perturbation detection and input preprocessing (denoising, compression-reconstruction) to attenuate adversarial features; Establish a second independent verification for high-risk interception so release is not decided by a single model