security · ai · llm

Prompt injection 101: a defender’s guide.

Ramswaroop
12 Sept 2026
securityaillm

SQL injection happened because code and data shared a channel. Prompt injection is the same mistake, reborn: to a language model, your instructions and a stranger's web page are both just text.

Yousummarise this pageAI agentreads + actsWeb pagehidden instructionsYour toolsemail · files"ignore the user…"
the page was supposed to be data. It filed itself under instructions.

Two flavours

Why it is genuinely hard

There is no reliable "this part is data" flag inside a model's input. You can ask nicely, you can add delimiters, you can use a second model as a filter — all of that helps, none of it is a guarantee. So design as if the model will sometimes be fooled.

The defences that actually earn their keep

  1. Least privilege. If the agent cannot send email, a hijacked agent cannot send email.
  2. Separate duties. The component that reads untrusted text should not also hold the dangerous tools or the secrets.
  3. Human approval for irreversible or outward-facing actions. Show exactly what will happen.
  4. Validate outputs and tool arguments like any other untrusted input.
  5. Limit exfiltration paths — no free-form URLs, no rendering arbitrary remote images from model output.
  6. Log and monitor. You cannot fix what you cannot replay.
ethics note

Test only systems you own or have permission to test. Responsible disclosure, always. The goal is to build things that survive the internet, not to ruin someone's afternoon.

The mindset shift: stop asking "how do I make the model never fall for it?" and start asking "what is the worst thing it could do if it did — and have I made that boring?"


© Ramswaroop Patelwritten as one plain .html file