Misleading Output from Explanation Tools
解釋工具產生誤導
Verification & Validation
Risk Description
When reviewers approve decisions based on feature importance explanations provided by an interpretability tool; due to unmitigated control gaps, it later emerges that the explanations are unstable and do not reflect the model's actual basis, requiring re-examination of every decision previously made on that footing, triggering compliance exposure and operational reputational costs.
Framework Mappings
EU AI ActArt.13
ISO/IEC 42001Annex A.6.2.4
NIST AI RMFMEASURE 2.9
ISO/IEC 5338驗證與確效
MIT AI Risk RepositoryDomain 7
Risk Treatment & Implementation Guidance
Validate the stability and faithfulness of explanation tools before adoption, confirming explanations reflect the model's actual basis; Cross-check with multiple explanation methods, never approving significant decisions on a single tool; Periodically audit explanation quality, suspending unstable tools and re-reviewing affected decisions