Skip to content

Research / source to practice

Read the finding.
Trace the boundary.

Three research starting points, connected to original lessons. Source observations, our interpretation, and sandbox limitations are kept distinct.

Application integration / untrusted content

Indirect prompt injection

Greshake et al. · 2023

Reported research
The researchers examined how instructions embedded in retrieved content can influence LLM-integrated applications. The attack’s reach depends on the application’s data access and available actions.
Boundary interpretation
Retrieved content is promoted into behavioural instructions while the assistant operates with application authority.
Defensive lesson
Treat document text as untrusted, constrain retrieval and downstream actions, and prove that a permission boundary holds even when the model proposes an unsafe operation.

Our tenant sandbox evaluates explicit retrieval authorization. It does not reproduce the paper’s target models or prove a natural-language injection.

Multimodal input / adversarial images

Image Hijacks

Bailey et al. · 2023

Reported research
This work studies adversarial images that steer the behaviour of vision-language models at runtime. Its attack surface is multimodal input; it is not a study of ordinary multi-turn text drift.
Boundary interpretation
Attacker-controlled visual input influences generated behaviour.
Defensive lesson
Define the expected image-processing task, evaluate adversarial input separately from ordinary images, and restrict consequential actions outside the model.

There is no executable adversarial-image lab in this release. The linked lesson teaches the general trust-boundary concept.

Model behaviour / generated jailbreaks

AutoDAN

Liu et al. · 2023

Reported research
AutoDAN explores automatically generated, semantically coherent jailbreak prompts using a genetic search approach. This is different from simply encoding text to evade a keyword filter.
Boundary interpretation
Generated user inputs attempt to redirect a model’s intended safety behaviour.
Defensive lesson
Evaluate varied inputs and repeated live runs. Keep model refusal testing distinct from application permission enforcement and record the model and evaluation conditions.

Our guided labs do not run AutoDAN or establish jailbreak resistance. They evaluate application controls on a finite deterministic suite.

Evidence before certainty.

A positive control, a model response, a committed action, and a reproduced vulnerability are different kinds of evidence. Our learning material names the observation and its limits.

Learn to write a reproducible result →