Lesson 2 / 6 · 2 minute read
Keep secrets out of reach
Understand the difference between prompt disclosure and exposure of protected information.
Prevent a support assistant from exposing a synthetic staff code while preserving useful answers.
- 01Customer question
- 02Context builder
- 03Model context
- 04Response
- 05Customer
Identify the protected value
A leaked system prompt is not automatically a high-impact vulnerability. Determine whether it exposes a credential, private record, proprietary information, or a capability that an attacker can use. In this sandbox the staff code is a synthetic protected value, not a real credential.
Minimize context
If a customer assistant does not need a staff code, do not put it in customer-facing model context. A stronger instruction to keep that code secret still leaves the value accessible to a component that generates text. Keep privileged operations and values behind application authorization.
Check utility
A defense that refuses every customer question avoids disclosure by abandoning the product’s task. Verify that the assistant still explains returns and opening hours. In a real model evaluation, add varied attack phrasing and repeat runs; a deterministic sandbox only demonstrates its configured rules.
Check your understanding
Which result best supports a working defense?
Lesson completion is a self-recorded learning milestone on this device. Lab results are tracked separately.
Put it into practice
Make the boundary observable.
Start with an explicit action in the guided sandbox. Then explore the related mission’s execution mode and evidence.
Sources and further reading
These sources inform the concepts. Our examples and sandbox scenarios are original and synthetic.
OWASP LLM Top 10 — 2025 taxonomy ↗