S0150RS6-T20-S0150-Z · Full risk code

Generation of Hate Content Targeting Groups

生成針對群體的仇恨內容

Operation & Monitoring
Risk Description

When an attacker gradually induces the system under the pretext of creative research, obtaining large volumes of denigrating content targeting a specific group and distributing it online; due to unmitigated control gaps, the organization's safeguards function correctly at the level of each individual request but fail to recognize the overall intent, triggering compliance exposure and operational reputational costs.

Framework Mappings

EU AI ActArt.5、Art.50
NIST AI 600-1Dangerous, Violent, or Hateful Content
NIST AI RMFMANAGE 2.4
ISO/IEC 5338運作與監控
MIT AI Risk RepositoryDomain 4

Risk Treatment & Implementation Guidance

Deploy harmful-content detection and layered output filtering covering group-directed disparagement; Apply cross-turn intent analysis to progressive elicitation rather than single-turn review; Act on confirmed abusive accounts and retain records supporting platform governance