Lesson 3 / 6 · 2 minute read
Follow an indirect instruction
Trace a malicious instruction from retrieved content to an attempted action.
Distinguish attacker content being retrieved from that content causing an unauthorized operation.
- 01Attacker document
- 02Retrieval
- 03Agent context
- 04Proposed tool call
- 05Recorded effect
Prove consumption first
A failed attack tells you little if the application never read the planted content. Use a harmless marker or an instrumented retrieval trace to show which document reached the model. Retrieval is a positive control, not proof that an instruction was obeyed.
Watch the next boundary
A malicious note might request a new task, ask for data from another tenant, or propose an external publication. The important transition is from untrusted content to a decision using application authority. Tool arguments and committed state provide stronger evidence than a model announcing success.
Constrain what the agent can do
Limit retrieval to the authenticated tenant and constrain every downstream tool independently. Keeping retrieved text in a separate field can improve model behaviour, but permission decisions should not depend on the model respecting that field’s label.
Check your understanding
A document marker appears in the answer. What has been established?
Lesson completion is a self-recorded learning milestone on this device. Lab results are tracked separately.
Put it into practice
Make the boundary observable.
Start with an explicit action in the guided sandbox. Then explore the related mission’s execution mode and evidence.
Sources and further reading
These sources inform the concepts. Our examples and sandbox scenarios are original and synthetic.
Greshake et al. — indirect prompt injection research ↗