Copyright Disputes Over Web-Scraped Content
網路爬取內容涉著作權爭議
When an organization trains a commercial model on a publicly scraped dataset; due to unmitigated control gaps, rights holders observe that model outputs closely resemble their works and assert infringement, triggering external stakeholder impacts and causing the organization can neither demonstrate the licensing status of its data sources nor readily remove specific content from the trained model.
Framework Mappings
Risk Treatment & Implementation Guidance
Conduct copyright review of training datasets, recording each source's license status and excluding content of unclear provenance; Manage training-data licensing so commercial models prefer clearly licensed data or recognized legal exceptions; Deploy similarity detection against known works at the output layer to reduce infringement exposure