S0042RS3-T08-S0042-Z · Full risk code

Manipulation of Human Preference Data

人類偏好資料遭操縱

Design & Development
Risk Description

When an organization's preference annotation process lacks monitoring of annotator behavior; due to unmitigated control gaps, some annotators consistently favor manipulative responses, and after tuning the model exhibits ingratiating and steering tendencies, triggering external stakeholder impacts and causing the organization notices the behavioral shift only after user reports.

Framework Mappings

OWASP Top 10 for LLMLLM04
MITRE ATLASAML.T0020
NIST AI 100-2Poisoning
EU AI ActArt.5
ISO/IEC 5338設計與開發
MIT AI Risk RepositoryDomain 2

Risk Treatment & Implementation Guidance

Apply multi-layer verification to preference labeling, with key samples independently labeled by multiple annotators; Deploy annotator behavior analytics detecting systematic bias toward particular response styles; Validate post-tuning behavior against baselines to confirm no sycophantic or manipulative drift