Application integration / untrusted content
Indirect prompt injection
Greshake et al. · 2023
- Reported research
- The researchers examined how instructions embedded in retrieved content can influence LLM-integrated applications. The attack’s reach depends on the application’s data access and available actions.
- Boundary interpretation
- Retrieved content is promoted into behavioural instructions while the assistant operates with application authority.
- Defensive lesson
- Treat document text as untrusted, constrain retrieval and downstream actions, and prove that a permission boundary holds even when the model proposes an unsafe operation.
Our tenant sandbox evaluates explicit retrieval authorization. It does not reproduce the paper’s target models or prove a natural-language injection.
Multimodal input / adversarial images
Image Hijacks
Bailey et al. · 2023
- Reported research
- This work studies adversarial images that steer the behaviour of vision-language models at runtime. Its attack surface is multimodal input; it is not a study of ordinary multi-turn text drift.
- Boundary interpretation
- Attacker-controlled visual input influences generated behaviour.
- Defensive lesson
- Define the expected image-processing task, evaluate adversarial input separately from ordinary images, and restrict consequential actions outside the model.
There is no executable adversarial-image lab in this release. The linked lesson teaches the general trust-boundary concept.
Model behaviour / generated jailbreaks
AutoDAN
Liu et al. · 2023
- Reported research
- AutoDAN explores automatically generated, semantically coherent jailbreak prompts using a genetic search approach. This is different from simply encoding text to evade a keyword filter.
- Boundary interpretation
- Generated user inputs attempt to redirect a model’s intended safety behaviour.
- Defensive lesson
- Evaluate varied inputs and repeated live runs. Keep model refusal testing distinct from application permission enforcement and record the model and evaluation conditions.
Our guided labs do not run AutoDAN or establish jailbreak resistance. They evaluate application controls on a finite deterministic suite.