S0202RS5-T17-S0202-Z · Full risk code

Fine-Tuning Breaking Safety Alignment

微調破壞安全對齊

Re-evaluation
Risk Description

When leaked model weights are fine-tuned by a third party with a small amount of data and circulated publicly with safeguards removed; due to unmitigated control gaps, although the original developing organization is not the actor, the identifiable provenance of the model leaves it facing reputational and liability disputes, triggering compliance exposure and operational reputational costs.

Framework Mappings

OWASP Top 10 for LLMLLM03
MITRE ATLASAML.T0018
NIST AI 100-2Poisoning
ISO/IEC 5338重新評估
MIT AI Risk RepositoryDomain 4

Risk Treatment & Implementation Guidance

Build alignment robustness in training to raise the cost of removing safeguards via light fine-tuning; Prepare contingency plans for weight leakage, including source-identifiable marking and public-statement procedures; Monitor circulating derivatives, initiating reporting and takedown requests when safeguard-stripped versions appear