S0029RS4-T16-S0029-Z · Full risk code

Synthetic Data Amplifying Existing Bias

合成資料放大既有偏誤

Design & Development
Risk Description

When an organization uses large volumes of generated data to expand its training set for a new model version; due to unmitigated control gaps, after several iterations, the model's handling of edge cases visibly deteriorates, outputs become homogeneous, and previously corrected biases reappear, triggering external stakeholder impacts and causing investigation traces the cause to distributional narrowing in the generated data.

Framework Mappings

EU AI ActArt.10
NIST AI 600-1Harmful Bias and Homogenization
NIST AI RMFMEASURE 2.11
ISO/IEC TR 24027
ISO/IEC 5338設計與開發
MIT AI Risk RepositoryDomain 1

Risk Treatment & Implementation Guidance

Test synthetic data for bias and distributional narrowing against the original; Keep a minimum proportion of original data so synthetic data does not dominate; Monitor diversity and edge-case capability per generation, adjusting composition on regression