S0058RS1-T03-S0058-Z · Full risk code

Misleading Output from Explanation Tools

解釋工具產生誤導

Verification & Validation
Risk Description

When reviewers approve decisions based on feature importance explanations provided by an interpretability tool; due to unmitigated control gaps, it later emerges that the explanations are unstable and do not reflect the model's actual basis, requiring re-examination of every decision previously made on that footing, triggering compliance exposure and operational reputational costs.

Framework Mappings

EU AI ActArt.13
ISO/IEC 42001Annex A.6.2.4
NIST AI RMFMEASURE 2.9
ISO/IEC 5338驗證與確效
MIT AI Risk RepositoryDomain 7

Risk Treatment & Implementation Guidance

Validate the stability and faithfulness of explanation tools before adoption, confirming explanations reflect the model's actual basis; Cross-check with multiple explanation methods, never approving significant decisions on a single tool; Periodically audit explanation quality, suspending unstable tools and re-reviewing affected decisions