Prompt injection is a supply-chain problem now
When your agent reads email, tickets and web pages, every one of those sources becomes part of your attack surface. We need to start treating it that way.
For the first couple of years of LLM apps, prompt injection was mostly a party trick. Someone typed "ignore previous instructions" into a chatbot, it said something silly, and a screenshot went around. The blast radius was one conversation.
That era is over. Agents now read inboxes, browse the web, open pull requests and call internal APIs. The attacker no longer needs to talk to your bot. They just need to put text somewhere your bot will eventually read.
The shift: from user input to every input
In traditional software security, we learned to distrust user input. In agentic systems, the line between "user input" and "data" disappears. A calendar invite, a product review, a PDF attached to a support ticket, or a code comment in a dependency can all carry instructions.
This is structurally the same problem as a software supply-chain attack. You're executing content you didn't write, from sources you don't control, with the permissions of your own system.
Every document your agent can read is code it might run.
What we're seeing in the wild
Across customer red-team engagements this year, indirect injection was the single most common critical finding. A few anonymized examples:
- A merchant descriptor in a banking app's transaction history told the dispute assistant to mark every claim as approved.
- A README file in a public repo instructed a coding agent to print environment variables into a PR description.
- A hidden line of white text in a résumé told an HR screening assistant to rank the candidate first.
- A support ticket asked a triage agent to forward the previous ten tickets to an external address.
None of these required access to the app. All of them worked on the first try against an unguarded system.
Defenses that actually help
Separate privileges by source
Content from untrusted sources should not be able to trigger high-risk tools. If a retrieved web page is in context, the agent shouldn't be able to send email in the same step without confirmation. This is the single highest-leverage control.
Inspect tool outputs, not just user messages
Most guardrails only look at what the user typed. Probe Guard inspects every tool result and retrieved chunk before it enters the context window, and flags instruction-like content from untrusted sources.
Log provenance
When something goes wrong, you need to know which source introduced the instruction. Probe Trace records the origin of every chunk in context so incident review takes minutes instead of days.
Test it continuously
Your sources change daily. Run indirect-injection attacks in CI against realistic fixtures: poisoned emails, documents, and API responses.
The takeaway
If your security model for AI is "we filter what users type", you're defending the front door of a building with no walls. Map every source your agents can read, rank them by trust, and make sure your guardrails sit on all of them.