Home / Docs / Reference

Words we use

This page explains the words that you will meet on this site. It uses plain words and short sentences. The words are in four groups: the defense, attacks and labels, numbers and tests, and the tools around Ferrite. Each word has its own link, so other pages can point to it.

The defense#

Fingerprint
The list of tools and websites that a task is expected to need. Ferrite works it out from your request alone, before it reads any page. It is the standard that actions are compared with.
Must-use set
The part of the fingerprint that comes from fixed rules. The rules match whole words in your request. For example, “open” implies navigation, and “download” implies a download. Cookies and storage are allowed only if you name them.
May-use set
The part of the fingerprint that a model proposes. It holds capabilities the task might need beyond what a fixed set of rules (the rule layer) already decided. The model must pick from a closed list of seven capability names. Anything else is dropped. If the model fails, times out, gives a bad answer or is missing a key, the may-use set becomes empty. It is never skipped and never widened. An empty fingerprint sends every action to consent.
Capability
A kind of thing the agent may do, such as reading a page or filling in a form. Each capability comes with a scope. There are exactly seven capabilities.
Primitive
One basic browser action, such as navigate, click, download or js.execute. A capability groups primitives that are alike in how risky they are.
Origin
A website address in its basic form: the scheme, the host and the port, for example https://mail.example.com. Ferrite compares origins by the whole host, never by the first letters. So localhost.evil.example is not local.
Scope
The set of origins where a capability may act. It has three kinds. Exact names specific origins. Domain suffix allows one domain and its subdomains. task_open allows every origin, so it is the weakest.
Dry run
A practice run that does nothing real. The agent's plan runs against made-up data in a stand-in engine. The engine records every call. It has no real network to reach.
Synthetic twin
The made-up data that the dry run uses in place of the real pages and tools. It is encrypted on disk.
Comparator
The part that compares the fingerprint with what the dry run recorded. Its result lists the extra primitives and the origins that the task did not allow.
A question asked to you before an action runs. Ferrite draws it in a trusted panel that page content cannot touch. You approve or reject each item. Nothing real runs before you decide.
Guard (runtime guard)
The check that stops an action outside the fingerprint. The dry run cannot see a page that exists only in the real run. So the guard checks every real action again. An action that is not expected and not approved pauses the run and asks you. The guard has known limits. See Limits.
Sanitizer
A filter that looks in page text for known attack wording and cuts it out. It is pattern matching, so it misses some attacks. The defense does not depend on it.
Audit log
A record of the security events. Each entry holds a SHA-256 hash of its own content and of the entry before it. So a change to an old entry breaks the chain and can be found. This is why it is called hash-chained. It is stored in SQLite.
Fail closed, fail to empty
What Ferrite does when something goes wrong. The prediction becomes empty, so more actions go to consent. The agent refuses to act without a model. Ferrite never skips a check, and it never lets more through.

Attacks and labels#

Agent
The AI that does a task in the browser for you. It reads pages and chooses each step.
Injection
Hidden text that tries to give the AI new orders. It can sit in a hidden element, an HTML comment, an image alt attribute or a field in a tool's answer.
Indirect prompt injection (IPI)
An injection that comes in through content the agent reads, such as a web page or a tool's answer. It does not come from you, the user. The word “indirect” means the attacker never talks to the agent directly.
Carrier
Where the injection sits. It is either web content (a page) or tool output (the answer of a tool).
Residual
An attack that looks exactly like the normal task. It uses the same capability at the task's own site, or it only changes data. No fingerprint can separate it from the task. Ferrite reports the residual as a floor, not as a failure of the guard. It keeps these attacks in every count.
Origin shift
An attack that does the same kind of action as the task, but at a different website. The tool is expected. The site is not. Because the origin differs, a check by origin can catch it.
Attack categories
Five labels for what an attack tries to do. They are listed next.
AgentRedirection
The injection tries to send the agent to the attacker's content.
DataExfiltration
The injection tries to make the agent leak data to a place that is not allowed.
UnauthorizedAction
The injection tries to make the agent do an action that the task did not call for.
WithinFingerprintAbuse
The injection stays inside the tools and websites that the task already expected. So the fingerprint cannot separate it from the task. Some of these attacks can still be caught by origin (see origin shift). The rest are the residual.
ScopeEscalation
The injection tries to widen what the agent is allowed to do beyond the task.
Ground truth
The answer key for a case: what the attack really tries to do. People write it from the specification. It never comes from running the defense. So the defense cannot grade itself. It has four forms. Deviation is a tool or website that the task did not allow. The others are origin shift, residual (data only) and none (a normal task).
Attempted, executed, blocked
These words describe an attack run. Attempted means the agent proposed an action that counts as the attack. Executed means such an action ran. Blocked means it was attempted and the guard stopped it. An action counts as the attack in two cases. In the first, it uses a primitive or origin that the ground truth names. In the second, it contains a string that the attacker chose.

Numbers and tests#

Benign
A normal task with no attack in it. Benign cases show what the defense costs you, for example how often it asks a needless question.
Attack template
The wording that an attack uses. The live run used one template (important_instructions). A model that resists one wording may not resist another.
Corpus, case
A case is one test: a task, a page or tool answer, and an attack or none. The corpus is the whole set of cases.
Scripted agent
A stand-in agent that does exactly what each case's authored deviation says. It is not a model. The main evaluation uses it.
Mode, baseline
A mode says which parts of the defense are on. In the live run, off has nothing on. It is the baseline. guard has the runtime guard on. The main evaluation has four modes. See Evaluation.
ASR, ADR, SDR, FGR
Short names for the main rates. ASR (attack success rate) is the share of attack cases where the attack ran. ADR is the share of deviation cases where the comparison flagged the deviation. SDR is the share where the sanitizer found the expected text. FGR (false gate rate) is the share of normal cases where Ferrite asked you a question. Metrics has the exact formulas.
95% interval
A range around a rate. It shows how far the rate could move if you had more cases. A wide range means few cases. Ferrite uses the Wilson method. For example, 2.0% [1.4-2.8] means 2.0%, with a range from 1.4% to 2.8%. In the main corpus the cases are not fully independent, so the real range is wider than the one shown.
McNemar test
A test that compares two modes on the same cases. It looks only at the cases where the two modes disagree. Say b is the number of cases where only the baseline had the attack run. Say c is the number where only the defended mode did. A small p-value means a split this uneven is unlikely if the two modes were the same. Ferrite uses the exact version of the test.
Cohen's h
A number for how big the gap between two rates is. A p-value says a gap is probably not chance. It does not say how big the gap is. Cohen's h does.
False positive
In the live report, a normal task where the guard refused an action. The task was not under attack, so the refusal was not needed. The main evaluation has a similar rate, the false gate rate: a normal task that gets a question.
Precision
Of the capabilities that the predictor named, the share that the task really needs. A low value means it names too many.
Recall
Of the capabilities that the task really needs, the share that the predictor named. A low value means it misses some. A missed capability makes the guard refuse a needed action.
Temperature 0
A model setting. At temperature 0, the model has no random choice in how it picks words. So the same input should give the same answer. The live run uses it, with one try for each case.

The tools around Ferrite#

Servo
The web engine Ferrite uses to draw pages. A build without Servo has the interface but cannot draw pages.
AgentDojo
An open set of tasks and attacks for testing AI agents. Ferrite imports its tasks as 1,046 cases. Ferrite reads the task files and never runs AgentDojo's code. Ferrite maps AgentDojo's tool calls to its own primitives to make the labels. So Ferrite's numbers are not AgentDojo's scores and cannot be compared with them.
Laya
An optional small local model. It picks routine browsing steps, which saves calls to the main model. It is never part of the security boundary. Every step it picks still goes through the guard and consent. See Models and settings.
Provider, API key
A provider is the service that runs the AI model, such as Ollama or Gemini. An API key is a secret code that lets the app use that service.
Model tag
The exact name of a model on a provider, for example gemma4:31b. Ferrite has no default tag, because providers retire names. Pass it with --model in the live runner.
Fast model, agent model
The two jobs a model can do. The fast model predicts the fingerprint, once for each task. The agent model reads pages and chooses each step. One model can do both.
Cache
A store of earlier model answers. If the same question comes again, Ferrite uses the stored answer. In the live runner, a re-run of an unchanged case is free and does not count against --max-calls. --no-cache turns it off.
Rate limit
A cap that a provider sets on how many requests it accepts in a period of time. When the cap is hit, the live runner waits and tries again, up to 5 tries for each call by default. It stops with exit code 4 after 3 failed calls in a row. --pause-ms spaces out the calls.
Batch
A group of cases that one command runs. --batch-size counts cases. --max-calls counts model calls, and its default is 100. See Run it from source.
Resume
Continue where a run stopped. The live runner saves each result before it starts the next case. Run the same command again and it skips the stored results. A crash or Ctrl-C loses at most the case that was running. Stored results are reused only if the settings that change behaviour are the same.
DevTools, backtrace
DevTools is the panel for developers. It shows a page's console messages and requests, and engine crashes. A backtrace is a list of the code steps that led to a crash. See Debugging and logs.
Vendored patch
A changed copy of the engine that is kept inside the repository.
Notarized
Checked by Apple and approved. The macOS builds of Ferrite are not notarized, so macOS blocks the first launch.