IA-06 (Safety and Security )
Ensures that the AI system, under defined conditions, does not cause harm or physical threat to human life, health, well-being, property, or the environment, while guarding against malicious threats such as adversarial attacks, data poisoning, and model theft.
Representative ScenariosS0132 Inference Resource Exhaustion AttackS0133 Hijacking of Computing ResourcesS0134 Crafted Inputs Amplifying Computational CostS0137 Cost Amplification AttackS0140 Autonomous System Accessing Files Beyond AuthorizationS0141 Unauthorized Outbound Connections by an Autonomous SystemS0178 Automated Cyber AttackS0182 Insufficient Security in Inter-System ProtocolsS0183 Tacit Coordination Among Autonomous SystemsS0184 Escalation Through System Interaction in Critical DomainsS0186 Multiple Systems Converging on Collective SuboptimumS0189 Tacit Coordination Through Shared Pricing Systems
Representative ScenariosS0036 Label Tampering AttackS0037 Backdoor Trigger Sample ImplantationS0039 Poisoning of the Retrieval Knowledge BaseS0040 Safety Alignment Broken During Fine-TuningS0041 Gradient Poisoning in Distributed TrainingS0043 Upstream Contamination of Automated Data PipelinesS0045 Malicious Packages Entering the Development EnvironmentS0046 Backdoored Models on Public PlatformsS0047 Vulnerabilities in the Training Framework ItselfS0048 Blind Spots from Missing Component InventoryS0049 Exploitation of Deep Learning Framework VulnerabilitiesS0050 Supply Chain Vulnerabilities in Preprocessing ToolsS0051 Insecure Serialization FormatS0070 Adversarial Perturbation of ImagesS0071 Adversarial Audio and Hidden CommandsS0072 Adversarial Text Evading DetectionS0073 Physical-World Adversarial AttackS0074 Cross-Modal MisdirectionS0075 Universal Adversarial PerturbationS0076 Cross-Model Transfer AttackS0077 Coordinated Multimodal PerturbationS0078 Functional Replication AttackS0088 Unencrypted Model Weights LeakedS0093 Inference Interface Without Access ControlS0094 Unpatched Container Image VulnerabilitiesS0095 Overly Permissive Cloud Service ConfigurationS0096 Inference Endpoint as an Internal PivotS0097 Multi-Tenant Isolation FailureS0118 Direct Prompt Injection Bypassing SafeguardsS0119 Indirect Prompt Injection Hijacking the SystemS0120 System Instruction DisclosureS0121 Multi-Step Inducement Attack ChainS0122 Tool Invocation Manipulated by InjectionS0123 Covert Goal HijackingS0124 Multimodal Hidden Instruction InjectionS0125 Cumulative Context PoisoningS0126 Retrieval Injection Hijacking OutputS0127 Capability Theft Through Query DistillationS0128 Side-Channel Inference of Model InformationS0129 Side-Channel Risk on Shared Computing PlatformsS0130 Residual Data in Shared Hardware MemoryS0150 Generation of Hate Content Targeting GroupsS0151 Generation of Virtual Content Harmful to MinorsS0153 Non-Consensual Synthetic Intimate ImageryS0154 Automated Generation of Harassment ContentS0155 Production of Extremist ContentS0156 Operational Guidance for Illegal ActivityS0157 Inappropriate Content and Guidance for MinorsS0158 Large-Scale Contamination by False InformationS0168 Synthetic Identity FraudS0169 Mass Production of False ContentS0170 Highly Customized Phishing AttackS0171 Lowered Barriers to Dangerous KnowledgeS0174 Misuse of Scientific Reasoning CapabilityS0175 Precision Social EngineeringS0176 Manipulation of Elections and Political ProcessesS0177 Distribution of Non-Consensual Synthetic Intimate Imagery
Representative ScenariosS0053 Replacement of Fine-Tuning Layer FilesS0055 Systemic Risk from Model HomogeneityS0089 Removal of Model MarkersS0090 Weight Tampering During TransmissionS0091 Model Substitution at DeploymentS0092 Hijacking of the Update Channel