S0040RS3-T08-S0040-Z · Full risk code

Safety Alignment Broken During Fine-Tuning

微調階段破壞安全對齊

Design & Development
Risk Description

When an organization outsources model fine-tuning without verifying the contents of the fine-tuning dataset; due to unmitigated control gaps, after deployment, the model produces content that should have been blocked under specific inducement, and only then does the organization discover that safety alignment was compromised during fine-tuning, triggering compliance exposure and operational reputational costs.

Framework Mappings

OWASP Top 10 for LLMLLM04
MITRE ATLASAML.T0018
NIST AI 100-2Poisoning
ISO/IEC 5338設計與開發
MIT AI Risk RepositoryDomain 2

Risk Treatment & Implementation Guidance

Security-review outsourced fine-tuning datasets for alignment-breaking samples; Run alignment verification after fine-tuning, comparing safety-behavior baselines before and after; Continuously monitor guardrail effectiveness post-deployment, detecting anomalous outputs under adversarial prompting