How the defense works
Ferrite's guarantee is not “this text is safe”. It is “this action was inside what your task needed, or you approved it”. This page walks through how that is decided, stage by stage. It also lists the invariants (rules that always hold) that keep it together.
Some words on this page need a short explanation. The full list is on Words we use.
- Fingerprint
- The list of tools and websites a task is expected to need.
- Guard
- The check that stops an action outside that list.
- Dry run
- A practice run that does nothing real.
- Consent gate
- A question asked to the user before the action runs.
- Sanitizer
- The text filter that cuts out text that looks like orders to the AI.
- Origin
- The address of a website: its scheme, host and port.
The capability model#
A capability is not an invented tool such as “email.read”. It is an action class. This is a group of primitives (basic actions) that share a security character. Each capability is paired with an origin scope, which says where it may act. “Email” is a property of the scope. The task's author sets it. It is never a tool the model can make up.
| Capability | Allows |
|---|---|
web.read | Reading page content. |
web.navigate | Going to a page. |
web.interact | Clicking, typing, filling and submitting. |
web.download | Downloading files. |
scoped.read | Cookies and web storage, only when named. |
clipboard.read, clipboard.write | Clipboard access. |
The vocabulary is closed. There are exactly seven capabilities. We filter a model's output against that list. A scope is one of three typed kinds. They run from tightest to loosest:
- exact
- Specific known origins. This is the tightest scope.
- domain_suffix
- A limited family, such as
*.wikipedia.org. - task_open
- Open browsing. It needs a written reason. It is reported as a weak scope, because under it every origin is admitted.
Each capability has its own scope. So a task can combine a narrow cookie read on one site with a wide page read for search results. Neither widens the other. We compare origins after normalizing them. We compare the whole host, never a prefix. localhost.evil.example is not local.
Invariant: js.execute is unconditionally unscopable
Running any JavaScript can create any other primitive past the tool boundary. So no capability may ever expect it, at any scope. It is always a deviation, and it always needs consent. The type system enforces this. No capability variant could authorize it. So a prediction cannot even express it.
1 · Predict#
Before any page is fetched or any tool runs, Ferrite turns your request alone into an expected fingerprint. It has two layers:
- The rule layer (
must_use) is fixed, offline and pure. It matches whole words in your request. “open” or “go to” means navigation. “summarize”, “list”, “search” and other everyday read verbs mean a page read. “download” means a download. “fill”, “submit” or “send email” mean interaction. Cookies or storage are granted only when you mention them. It does not match parts of words. So “information” does not grant form filling. An open-ended prompt that matches nothing gives an empty fingerprint. An empty fingerprint sends everything through consent. - The model layer (
may_use) asks the fast model you set up what the task might need beyond the rules. The answer is untrusted text. Each label must be exactly one of the seven capability names. Any other label is dropped at the boundary.
Fail to empty, never to a bypass. The model layer becomes the empty set after a provider error, a timeout, a rate limit, a malformed or empty body, or a missing key. It never skips fingerprinting. It never widens what is admitted. An empty prediction sends everything through consent. An outage costs you more prompts, not less safety.
2 · Sanitize#
Before content reaches the agent, a sanitizer looks for text that looks like orders to the AI, and cuts it out. It lowers risk. It is not the foundation. The architecture still holds with the sanitizer switched off. That is how the evaluation measures the two apart.
- Matching on meaning, not spelling. The sanitizer matches text in cleaned-up views. In these views, zero-width characters are removed. Look-alike and full-width letters are folded. Escapes (HTML, URL,
\u) are decoded. Leetspeak, ROT13, reversed text, base64 and hex are decoded. Invisible Unicode tag characters are handled. Each byte maps back to the original span. So detection reports the original snippet, and excision cuts the original sentence. - Pattern set v2 covers instruction override, agent-addressing, concealment, exfiltration (sending data out), chat-template and action-mimicry patterns, markdown-image exfiltration and prompt extraction. It works in eight languages. It also has a hidden-Unicode check.
- Everywhere a digest shows text. The scan covers attribute values (
alt,title,placeholder,aria-label,href), CSS-generated text and the raw page text. It does not cover only the visible body. - Linear time. Excision takes time linear in the input size. A 1 MB page with no sentence terminators is handled without a hang.
It is still pattern matching. By design, it misses paraphrase, languages outside the list, spelled-out markup, reversed non-Latin text and payloads split across elements. The test suite pins them as known misses. So the list stays honest.
3 · Dry-run#
The agent's plan runs against a synthetic twin. This is fake but realistic data. A stand-in engine serves it. The engine records every call and does no real I/O. Containment here comes by construction. It does not need an interceptor that could be bypassed. The dry-run engine has no real network to reach.
- Twin data is encrypted at rest. The key comes from the OS keyring or an environment variable. It is never a constant compiled into the binary.
- A download or a new tab is recorded at its destination origin, and as a network attempt. Suppose it were recorded at the page's own origin. Then an agent on an allowed site could fetch from an attacker without making a deviating event.
- A whole-turn timeout returns the partial record. It does not throw the record away.
4 · Compare#
The comparator takes the expected fingerprint and the dry-run record. It makes a diff. The diff lists extra primitives (capabilities the task did not imply) and out-of-scope origins (places the task did not authorize). One function, classify_event, classifies every event. The comparator and the runtime guard share it. So they cannot disagree. A sweep test covers every combination of primitive, origin and fingerprint.
5 · Consent#
If the diff is not empty, nothing real runs. The review panel (titled Review before running) lists each deviation in plain words. It asks you about each item: approve or reject. The application draws it, outside page content. Only the items you approve are added to what the real run may do. Each line names the offending action and the origin your request actually authorized. So a look-alike domain shows up as one. Proceed with approved stays disabled until every item is decided. So there is no way to approve by default.
The cards are calm on purpose. They have a neutral surface, a coloured bar down the left edge (amber undecided, green approved, red rejected), and soft tinted buttons. The buttons are not solid red and green. A test checks that the button labels stay readable in both themes.
The runtime guard#
The dry run runs against synthetic pages. So a page that carries an injection exists only in the real run. The dry run cannot see it. Suppose the loop compared once and then ran every real action. Then a deviation that only a real page caused would be compared against nothing. The guard closes that gap.
- Every real action is classified. It does not run unless it is in expected ∪ approved. That set is what the predicted capabilities and their scopes admit, plus what you approved in this task's consent prompt.
- The guard checks the effect of an action. A URL action is checked at that URL's origin. Anything else is checked at the active tab's origin. A click by
@refon a link whosehrefleads elsewhere counts as a navigation there. A click on a link to the site the tab is already on counts as either a click or a navigation. So a task that was expected only to navigate (“go to this site's docs page”) can follow the site's own links. A link to any other site, or a button, still needs its own permission. - An action that is neither expected nor already approved pauses the run and asks you. You can choose Allow once, Allow for task or Don’t allow. (Esc means “no”.) The question names the action and the site. It never shows text the page wrote. Nothing outside the prediction runs until you say yes.
- A block gives the agent a fixed sentence. The sentence has nothing from the action or the page. An observation goes back to the model. Anything the attacker chose, if echoed there, is itself an injection channel.
- The block is shown to you and written to the audit log (
CapabilityDenied). The fourth block ends the run. - An empty prediction blocks everything that has no consent.
The guard has stated limits. It cannot work out the destination before the action runs in three cases. These are a click by CSS selector, a form action (submit, Enter in a field) and a server-side redirect. So the guard checks the tab's new origin on the next action, after the request was made. See Limits.
With a real model, the guard stopped 25 of the 31 attacks that ran without it. It stopped none of the 5 AgentRedirection attacks. See Evaluation: Live results.
The audit log#
Ferrite appends security events to a SQLite-backed log. Every entry commits to the one before it:
- Hash chain. Each entry's hash covers its content and the previous entry's hash. It also has a fixed domain-separation prefix and a schema version. So reordering or editing an entry breaks every hash after it.
- Full coverage. Every stored column is in the hash input. An earlier version left out two columns. They could be rewritten while verification still passed. That was fixed. A tamper test matrix covers it.
- Canonical form. The hash input is written as JSON in a fixed field order. It is not
Debugoutput. Sonulland the empty string can be told apart. Every string is length-delimited. - Verifiable.
verify_chain()recomputes the chain. The evaluation records one entry per run (96 in the original run) and checks it.
Two logs, on purpose
The model-activity trace (every prompt, answer and timing) is kept apart from the audit log. The audit log is small, hash-chained and about decisions. The trace is big and local. Mixing them would make the log neither small nor tamper-evident in a useful way.
Invariants at a glance#
- The fingerprint is built only from text you wrote. Page text, agent output and steps never feed it.
- A model's labels pass a closed list. Anything else is dropped.
- Any provider failure empties the prediction. It never bypasses it.
js.executecan never be expected, at any scope.- The comparator and the guard classify the same way, because they share one function.
- The agent never types into a password, one-time-code or card field. The page script refuses. A page that needs one pauses the run for you (the sign-in handoff).
- No test needs the network or an API key. Live-provider tests are opt-in.