Skip to content
← Learning paths

Lesson 2 / 6 · 2 minute read

Keep secrets out of reach

Understand the difference between prompt disclosure and exposure of protected information.

By the end

Prevent a support assistant from exposing a synthetic staff code while preserving useful answers.

Follow the boundary
  1. 01Customer question
  2. 02Context builder
  3. 03Model context
  4. 04Response
  5. 05Customer

Identify the protected value

A leaked system prompt is not automatically a high-impact vulnerability. Determine whether it exposes a credential, private record, proprietary information, or a capability that an attacker can use. In this sandbox the staff code is a synthetic protected value, not a real credential.

Describe demonstrated impact rather than treating all prompt disclosure as equivalent.

Minimize context

If a customer assistant does not need a staff code, do not put it in customer-facing model context. A stronger instruction to keep that code secret still leaves the value accessible to a component that generates text. Keep privileged operations and values behind application authorization.

Removing unnecessary sensitive context reduces the available disclosure surface.

Check utility

A defense that refuses every customer question avoids disclosure by abandoning the product’s task. Verify that the assistant still explains returns and opening hours. In a real model evaluation, add varied attack phrasing and repeat runs; a deterministic sandbox only demonstrates its configured rules.

Security and useful task completion are separate requirements.

Check your understanding

Which result best supports a working defense?

Choose one answer

Lesson completion is a self-recorded learning milestone on this device. Lab results are tracked separately.

Put it into practice

Make the boundary observable.

Start with an explicit action in the guided sandbox. Then explore the related mission’s execution mode and evidence.

Sources and further reading

These sources inform the concepts. Our examples and sandbox scenarios are original and synthetic.

OWASP LLM Top 10 — 2025 taxonomy ↗