security · ai · llm
Prompt injection 101: a defender’s guide.
SQL injection happened because code and data shared a channel. Prompt injection is the same mistake, reborn: to a language model, your instructions and a stranger's web page are both just text.
Two flavours
- Direct — the user types something that tries to override the rules.
- Indirect — the hostile text hides in content the model reads: a web page, an email, a PDF, a repo issue. The user never sees it.
Why it is genuinely hard
There is no reliable "this part is data" flag inside a model's input. You can ask nicely, you can add delimiters, you can use a second model as a filter — all of that helps, none of it is a guarantee. So design as if the model will sometimes be fooled.
The defences that actually earn their keep
- Least privilege. If the agent cannot send email, a hijacked agent cannot send email.
- Separate duties. The component that reads untrusted text should not also hold the dangerous tools or the secrets.
- Human approval for irreversible or outward-facing actions. Show exactly what will happen.
- Validate outputs and tool arguments like any other untrusted input.
- Limit exfiltration paths — no free-form URLs, no rendering arbitrary remote images from model output.
- Log and monitor. You cannot fix what you cannot replay.
Test only systems you own or have permission to test. Responsible disclosure, always. The goal is to build things that survive the internet, not to ruin someone's afternoon.
The mindset shift: stop asking "how do I make the model never fall for it?" and start asking "what is the worst thing it could do if it did — and have I made that boring?"