Upstream Contamination of Automated Data Pipelines
自動化資料管線上游污染
When an organization's data pipeline automatically ingests from multiple public sources without content review; due to unmitigated control gaps, after an upstream source is contaminated, harmful content enters the training set directly, triggering external stakeholder impacts and causing because the pipeline is highly automated, the problem spans multiple model versions and is extremely difficult to trace and remediate.
Framework Mappings
Risk Treatment & Implementation Guidance
Assign trust ratings to automated-pipeline sources, requiring content review before low-trust data enters training sets; Install automated quality gates in the pipeline intercepting harmful content and statistical anomalies; Snapshot source composition per training-set version so contamination can be traced to affected versions