Skip to content
← Learning paths

Lesson 1 / 6 · 2 minute read

Draw the trust boundary

Separate what the user authorized from what the model reads and proposes.

By the end

Identify the attacker-controlled input, the authorized task, and the component that must enforce permission.

Follow the boundary
  1. 01Authorized task
  2. 02Untrusted content
  3. 03Model proposal
  4. 04Permission check
  5. 05Action

Begin with the legitimate task

A support assistant may explain return policies. That task does not authorize revealing staff-only values, exporting other customers’ records, or issuing arbitrary refunds. Write down the permitted action and resource before considering an attack.

Authority comes from the authenticated task and application policy.

Content can influence a proposal

User messages, retrieved documents, tool descriptions, and remembered preferences can all affect model output. They have different origins. A convincing instruction inside a document remains document content; it cannot grant permissions to the agent reading it.

Track who controls each input and when the application consumes it.

Enforce at the consequence

Models are trained to follow instruction hierarchies, but model behaviour alone is insufficient for enforcing authorization. The tool, database, renderer, or approval service must check the operation against the task’s authority. A refusal is useful behaviour; a permission check is an enforceable boundary.

Inspect the component that can actually change state or release data.

Check your understanding

A retrieved document says a refund exception was approved. Where should the refund limit be enforced?

Choose one answer

Lesson completion is a self-recorded learning milestone on this device. Lab results are tracked separately.

Put it into practice

Make the boundary observable.

Start with an explicit action in the guided sandbox. Then explore the related mission’s execution mode and evidence.

Sources and further reading

These sources inform the concepts. Our examples and sandbox scenarios are original and synthetic.

OWASP LLM Top 10 — 2025 taxonomy ↗Greshake et al. — indirect prompt injection research ↗