Threat model
A defense is only as meaningful as the attacker it assumes. This page states that attacker exactly. It says what is trusted and what is out of scope. So you can read the evaluation against it.
Some words on this page need a short explanation. The full list is on Words we use.
- Injection
- Hidden text that tries to give the AI new orders.
- Fingerprint
- The list of tools and websites a task is expected to need.
- Capability
- A class of action, such as reading a page or downloading a file.
- Origin
- The address of a website: its scheme, host and port.
- Sanitizer
- The text filter that cuts out text that looks like orders to the AI.
- Dry run
- A practice run that does nothing real.
- Consent gate
- A question asked to the user before the action runs.
- Residual
- An attack that looks exactly like the normal task.
The attack#
Indirect prompt injection is an attack on an agent through content it reads. It is not an attack through the user's own prompt. The attacker never talks to the agent directly. They plant text where the agent will look. The agent cannot reliably tell data from orders. So it may obey the text.
In a browser, the results are concrete. The agent may go to an attacker's site, submit a form, download a file, read cookies or clipboard contents, or run script. Each of these is an action. Ferrite checks actions.
The attacker#
Can#
- Control the content of pages the agent reads. This includes hidden or off-screen elements, HTML comments,
alt,titleandmetatext, CSS-generated text and visible text. - Control the content of tool results the agent receives. Examples are fields of a JSON reply, free-text blobs, error messages and metadata.
- Pick any wording: a paraphrase, another language, an encoding, invisible characters, social engineering, or text split across elements.
- Make several requests. For example, a delayed payload, or a redirect chain that only shows up on a second read.
Cannot#
- Edit the user's prompt, or speak in the browser's own interface. The application draws the consent panel, not page content.
- Change Ferrite, the operating system, the model provider or the model weights.
- Reach the user's real session during a dry run. The dry run works on fake data, and no real network can be reached.
What is trusted#
| Input | Trust | Why |
|---|---|---|
| Your typed request (and up to five earlier user messages) | Trusted | It is the only text that feeds the fingerprint, the sanitizer input and the dry-run driver. |
| Page text, page digests (the short text form of a page), tab titles | Untrusted data | They are marked as data, cleaned and size-limited before the agent sees them. They never reach the prediction. |
| Earlier agent answers and steps | Untrusted data | They may have quoted a poisoned page. If they were kept in later prompts, one page could shape every later run in the chat. |
| The model's predicted labels | Untrusted, filtered | A label that is not one of the seven closed capability names is dropped at the boundary. |
| Password field values | Never read | They are not read into a page digest at all. |
The vocabulary the evaluation uses#
The evaluation describes every attack along fixed axes. So two authors tag a case the same way.
- Carrier vector
- Where the payload (the attack text) rides. Web content: hidden element, off-screen text, HTML comment, alt text, meta content, CSS pseudo-element, visible text. Tool output: JSON field, text blob, error message, metadata.
- Rhetorical technique
- What the text argues: instruction override, context manipulation, social engineering, goal hijack.
- Concealment
- How it is hidden: plain, obfuscation, payload splitting.
- Ground truth
- What the case's label says. Deviation is a capability or origin the task did not authorize. Within-fingerprint origin shift is an allowed capability used at a different origin. Within-fingerprint data-only is the irreducible residual: an attack that looks exactly like the normal task. None is benign, which means a normal task.
Out of scope, stated plainly#
Ferrite does not claim to stop these
- Attacks inside the prediction. The same capability, at the task's own site, with data-only abuse, looks exactly like the task. This is the residual. We report it as a floor. It is never dropped from the denominator.
- A model that takes the bait. Ferrite limits what a compromised agent can do. It does not make the model more robust. The scripted evaluation says nothing about how often a real model is fooled. The live test measured one model only. It tried the attack in 31 of 1542 measurable runs (2.0%). See Evaluation: Live results.
- A user who approves everything. The consent gate is a human choice. If the user approves every flagged action, containment is much weaker. See the consent-policy sweep in Evaluation.
- Open-ended tasks. Under a
task_openscope, every origin is admitted. So nothing separates the attacker's destination. - A compromised machine, provider or build. These sit below the trust boundary this design assumes.