The dangerous prompt injection is not the user typing a clever jailbreak. It is the instruction hidden inside the content your agent reads: a web page, a support ticket, a file, the output of a tool. The agent treats that content as if it were your instructions.
This is for security leaders who have signed off on AI agents and want to understand the attack that does not go through a login, a firewall, or a person.
Direct versus indirect
Direct prompt injection is someone typing "ignore your instructions" into a chat box. It is noisy and mostly a nuisance. Indirect injection is different: the malicious instruction rides inside data the agent consumes while doing its job. The agent fetches a page to research a task, and the page contains hidden text telling it to do something else. It reads a Jira ticket, and the ticket body carries a command. It gets a result back from a tool, and the result is crafted to redirect it.
The person never sees it. The agent does, and to the agent it looks like part of the task.
Why it is worse for agents than for chatbots
A chatbot that gets injected says something wrong. An agent that gets injected can act. It can run a command, write a file, call an API, move data, open a connection. The injection stops being a bad answer and becomes a real action on your systems.
The supply-chain framing
Here is the shift worth internalizing: every piece of content and every tool output an agent consumes is now part of your attack surface. It is a dependency you did not vet. You would not run untrusted code without controls; an agent effectively runs untrusted instructions every time it reads an unfamiliar page or an external result. The content is the supply chain, and most of it is unsigned.
How you actually defend against it
You cannot fully prevent injection at the model, and you should assume some attempts will get through. So the defense is layered, and the last layer is the important one: treat every fetched page and every tool output as untrusted, look for injection patterns in what goes into and comes out of the agent, and be able to stop the resulting action before it executes. Detection without a way to block is just a nicer incident report.
KonaSense does exactly this on the agent path. We scan the prompts and tool inputs an agent works with for injection and manipulation patterns, and because we sit at the point where the agent reaches out to act, we can deny the resulting action before the tool runs. The attempt becomes a blocked event with a record, not a successful action you discover later.
It is the same lesson as the rest of this series, in what an agent actually is: the risk is in the loop, at the moment of execution. Injection is one more way to reach that moment, which is exactly why the moment is where the control has to live.
See how KonaSense stops a hijacked action before it runs


