S0033RS3-T07-S0033-Z · Full risk code

Copyright Disputes Over Web-Scraped Content

網路爬取內容涉著作權爭議

Design & Development
Risk Description

When an organization trains a commercial model on a publicly scraped dataset; due to unmitigated control gaps, rights holders observe that model outputs closely resemble their works and assert infringement, triggering external stakeholder impacts and causing the organization can neither demonstrate the licensing status of its data sources nor readily remove specific content from the trained model.

Framework Mappings

EU AI ActArt.53
NIST AI 600-1Intellectual Property
ISO/IEC 42001Annex A.7.5
ISO/IEC 5338設計與開發
MIT AI Risk RepositoryDomain 6

Risk Treatment & Implementation Guidance

Conduct copyright review of training datasets, recording each source's license status and excluding content of unclear provenance; Manage training-data licensing so commercial models prefer clearly licensed data or recognized legal exceptions; Deploy similarity detection against known works at the output layer to reduce infringement exposure