S0120RS3-T10-S0120-Z · Full risk code

System Instruction Disclosure

系統指令外洩

Verification & Validation
Risk Description

When after obtaining the system instructions, an attacker designs inputs that precisely circumvent each protective rule; due to unmitigated control gaps, the organization's defensive design is largely nullified by the exposure of its details, requiring a complete redesign, triggering compliance exposure and operational reputational costs.

Framework Mappings

OWASP Top 10 for LLMLLM07
MITRE ATLASAML.T0051
NIST AI 600-1Information Security
ISO/IEC 5338運作與監控
MIT AI Risk RepositoryDomain 2

Risk Treatment & Implementation Guidance

Design against system-prompt disclosure, refusing to reveal own instructions and verifying regularly with elicitation tests; Keep sensitive rules and secrets out of the prompt, enforcing them in an external policy engine to limit disclosure impact; Design protections assuming eventual prompt exposure so security does not depend on prompt secrecy