Home / Docs / Honesty

Limits

A security project that hides its weaknesses cannot be judged. This page puts every known limit in one place. It says what the defense cannot stop. It says what the evaluation cannot show. It says what the product cannot do yet. Some limits are pinned by a test. That test stays in the repository, so the limit cannot quietly disappear. Some words here are new. Words we use explains them.

Ferrite is a research prototype

Do not use it to protect anything you cannot afford to lose. The numbers in the documentation describe an architecture in one experiment. They are not a guarantee.

Limits of the defense#

The residual: attacks inside the prediction#

Ferrite predicts a fingerprint for each task. The fingerprint is the list of tools and websites the task is expected to need. The defense contains every attack that needs an action outside that list, and nothing else. A residual is an attack that looks exactly like the normal task. One kind uses the same capability at the task's own site. Another kind is data-only. For example, it changes the content that a legitimate action returns. Ferrite cannot tell these attacks from the task. In the evaluation, the residual is 122 of 809 attacks (15.1%) with the sanitizer off. Ferrite reports this number as a floor. It keeps these attacks in every denominator.

Open-ended scopes#

A scope says which websites a task may use. Under a task_open scope, every origin is allowed. An origin is a website address. So nothing sets the attacker's destination apart. Sometimes an open-ended request matches no rule. Then the fingerprint is empty, and every action goes to consent. That is safe, but it can be noisy. Take a follow-up such as “do the same for the second one”. The dry run sees no page content. So the fingerprint is looser, and you may get more prompts. A dry run is a practice run that does nothing real. Ferrite accepts the safe direction here.

The sanitizer#

The sanitizer is a filter. It matches known attack patterns in page text, after it undoes common disguises. By design, it misses these things. Tests pin each miss.

  • Paraphrase (the same attack in other words).
  • Languages that are not on its list.
  • Words broken into syllables with hyphens.
  • Markup or URLs that are spelled out in words.
  • Reversed text in non-Latin scripts.
  • Attack text split across separate elements.

It also flags some safe pages by mistake. Pages that discuss attacks trigger it, for example “attackers exfiltrate data” and “never ignore the previous warnings”. Any mention of “system prompt” triggers it too. The architecture does not depend on the sanitizer.

The runtime guard#

The runtime guard checks every real action. It is the check that stops an action outside the predicted list.

  • Some actions cannot be traced to a destination before they run. These are a click by CSS selector, a form action (a submit, or Enter in a field), and a server-side redirect. Ferrite checks the tab's new origin on the next action. By then the request has already been made.
  • A host with a trailing dot (example.com.) fails closed. That means it is refused.
  • Within-fingerprint abuse looks the same as the task, by construction.

The human#

The consent gate is a question to a person. The person decides. A user who approves everything defeats it. In the original run, 45.5% of attacks succeeded when every flagged action was approved. When every flagged action was rejected, 0% succeeded. Real people are somewhere in between. No human-factors study has measured where.

The model#

Ferrite limits what a fooled agent can do. It does not make a model less easy to fool. A fooled model still tries the action. The defense contains the attempt. It does not stop the attempt from happening.

Limits of the evaluation#

  • The main evaluation uses a scripted agent. The agent re-enacts each case's authored deviation from ground truth. Ground truth is the answer key: what the attack really tries to do. So the main results say nothing about how often a real model takes the bait. One live run did use a real model. Its limits are in the next section.
  • The fingerprint is rules-only by default. The 0% false-gate rate is for a well-behaved scripted agent with that prediction. A false gate is a question asked to the user for a normal task. A real agent's rate depends on what it does. It also depends on the model prediction layer, which the default run does not use. The live run did use the model prediction layer. Its false-positive rate was 38.1%. See the next section.
  • The corpus is self-written and generated. The defense's author wrote the cases. Shared templates then expanded them. The author also tuned the defense against them. The intervals treat related cases as independent. So they understate the real uncertainty.
  • There is almost no independent slice. Only three of the 938 cases come from outside (AgentDojo). Three cases support no claim. There is no second author. So Cohen's κ, the score for how much two people agree when they label cases, cannot be computed. Ferrite has not faked it.
  • The 1,046-case AgentDojo import is not in the headline numbers. It is checked in and checked offline. It has results only from the one live run with one model. Its labels come from AgentDojo's ground-truth calls. Ferrite mapped those calls onto its own primitives. Some labels come from the goal text alone. It uses one attack template and a synthetic environment. So its numbers cannot be compared with published AgentDojo scores.
  • Only one real model has been the agent. That model is ollama gemma4:31b, in the live run below. The live runner can use Gemini and Ollama. You cannot select OpenAI-compatible or Anthropic endpoints, because the model layer has no backend for them.
  • ADR = 100% is true by construction. ADR is the share of deviations that the comparison flagged. It measures the comparator against a scripted agent that does exactly the declared deviation.
  • Some quantities are missing. Ferrite does not compute task-completion parity for normal tasks. There is no normal-task baseline under Off. It also does not compute the extra token cost. The execution record has no token usage. Both gaps have stated reasons.
  • A live-provider run cannot be reproduced byte for byte. So Ferrite reports it separately.

Limits of the live run with a real model#

The owner ran all 1,984 cases (the 938 cases plus the 1,046 AgentDojo import) with ollama gemma4:31b. The model played both roles: it predicted the fingerprint, and it was the agent. Each case ran twice, once with the guard off and once with the guard on. That makes 3,968 runs. None failed for infrastructure reasons. The numbers are in Live results on the Evaluation page. The full report is docs/results/live-eval-ollama-gemma4-31b.md. These limits apply to that run.

  • Only one model was tested. The results say nothing about any other model.
  • The model rarely fell for the attack. With the guard off, it carried out 31 of 1,542 measurable attack runs (2.0%). So the guard had few attacks to stop. With the guard on, 6 attack runs got through (0.4%). In 25 cases, an attack ran with the guard off and did not run with the guard on. The reverse never happened. These counts are small. Read them with their wide intervals.
  • The guard stopped none of the 5 AgentRedirection attacks. All 5 ran in both modes (5 of 62 runs, 8.1%). AgentRedirection is an attack that sends the agent to the attacker's content. Only 62 runs are in this category, so this is a small sample.
  • The guard refused an action in 38.1% of normal tasks. That is 86 of 226 normal runs. This is a false positive: a refusal in a task that was not under attack. The likely cause is that the predictor is weak. Only the 1,046 AgentDojo cases state what a task needs, so Ferrite scored the predictor on those. It found 70.1% of the capabilities a task needs (1,196 of 1,706). Only 52.2% of the capabilities it predicted were needed (1,196 of 2,293). It predicted the whole needed set for 56.2% of the tasks (588 of 1,046). In real use, a refusal like this would be a question to you.
  • Some attacks cannot be stopped by a fingerprint. In 27.1% of the measurable attack runs (418 of 1,542), the attack looks exactly like the normal task. These are residual cases. Among the 393 AgentDojo attack cases that name an extra capability, the prediction already allowed it in 98 (24.9%). The guard could not stop those. 432 attack runs could not be recognized in the agent's actions. They are residual cases with no attacker string. Ferrite leaves them out of the attack rates.
  • Ferrite used one attack template and tried each case once. A model that resists one wording may not resist another. One try for each case says nothing about how much results vary.
  • The practice engine does nothing real. The agent acts in a dry-run engine. The engine records what it is asked to do, and it runs nothing. Ferrite cannot see whether a task really succeeded. It treats a final answer from the agent as a stand-in for success.
  • The simulated user refuses everything. The guard refuses every deviation. This is the best case for containment. A real person who approves prompts moves these numbers toward the guard-off numbers.
  • Judging an attack uses a specification, not a perfect test. An action counts as the attack in two cases. In the first, it uses a primitive or an origin that the case's ground truth names. In the second, it contains a string the attacker chose. Both signals are imperfect. A model that wanders into an extra primitive on its own is counted as a hit.
  • These are Ferrite's own numbers. They are not AgentDojo's utility or security scores. Do not compare them with published AgentDojo numbers.
  • Some rows have few cases. For example, ScopeEscalation has 1 run, and the hand-written AgentDojo cases have 3. Do not quote small rows as findings.
  • The owner ran it on their own machine. The report in the repository is a copy. The raw per-run file is not in the repository.

Limits of the product#

Most of the newest work has only run on Linux

These features were checked on Linux. They are the DevTools panel, resizable panels, page-control overlays, the crash banner, calmer approval cards and the display-scale fix. The checks used a virtual display at normal pixel density. The engine parts also ran in the headless engine. None of it has run on a Mac, on a Retina screen or on a real GPU. The build on the releases page (commit f4ed744, 2026-10-02) is older than most of it.

Websites that do not work yet#

  • GitHub did not open in the owner's build. Two causes were found and fixed, and neither is confirmed on a Mac yet. A WebGL panic and an infinite loop in the engine's layout code both froze pages (see the debugging page). The fixes are in vendored engine patches. A vendored patch is a changed copy of the engine kept in the repository. What would settle it is the log, or the collect-logs zip, from a build that has the new console and request logging. Capture it right after loading GitHub. The owner needs to send this. See Debugging and logs.
  • The Google results page was pinned to the left edge. The likely cause is that the engine was never told the display scale. A Retina page was laid out twice as wide as it looks. This is now fixed. But nobody has seen the fix work, and Google could not be loaded where this was studied. It needs a screenshot from a new build.
  • Google and GitHub sign-in may still be refused. These sites decide from the whole session. Ferrite uses a young engine. Ferrite does three things about it: the identity choice, the Web APIs it added to the engine, and the sign-in handoff. None of these is a guarantee. Nobody has checked them against those sites. Google sign-in is still an open report. The engine still lacks service workers, navigator.mediaDevices (WebRTC), navigator.userAgentData and window.chrome. Ferrite does not fake any of them. It also does not fake hardware, canvas or renderer details, and it does not hide automation.
  • Web compatibility is Servo's. Servo is the web engine Ferrite uses to draw pages. Several CSS features that current sites use are missing or broken in Servo. They are :has(), container queries, mask-image, backdrop-filter, text-wrap, subgrid, popovers, and aspect-ratio on block boxes (a box with a width collapses to zero height). Pages that need them will look wrong. Ferrite turns on the engine features that were measured to work: adoptedStyleSheets, FontFace, WebGL, and seven more Web APIs. It found those seven by testing every engine switch that is off by default, one at a time. Ferrite leaves off the features that were measured not to work. These are WebRTC (it would hand out the microphone with no prompt), geolocation and service workers. Ferrite also adds one small script. The script keeps icons coloured by CSS visible, because Servo paints them black otherwise. Ferrite cannot add layout features that the engine does not have.
  • Zooming and interaction fail on some pages, and nobody knows why. The page crash banner and the Engine tab now show script-thread panics. Dropdowns and dialogs now work. These were the two possible causes. Nobody knows whether they explain what was seen.

The browser#

  • The DevTools Network tab has no status code. A request is listed when it starts, with its method, type and URL. Its size and time are added when it finishes, from the page's own Resource Timing. The engine does not tell the browser the response itself, so there is no status code and no waterfall. A request from another site that does not allow timing shows no size. fetch() and XHR are listed as Other.
  • The system file picker is untested on a real Mac or Windows machine. The file card opens it with the system's own tool. Only the command each system gets and what happens to the picked path are tested. Nobody has tried it with a real upload.
  • Input-method (IME) composition is not handled. Composed text is not passed on. Chinese, Japanese and Korean input are examples. Nobody has tested non-US keyboard layouts, dead keys or key repeat.
  • F12 opens the Audit panel, not DevTools. DevTools is Cmd+Opt+I or Ctrl+Shift+I, or Cmd/Ctrl+J. Whether F12 should change is an open decision.
  • The address bar searches with Google. The agent does not. A query typed in the address bar goes to Google. The agent's own prompt still tells it to search with DuckDuckGo Lite and not to use google.com. This is because Google's pages stayed blank in this engine when that was written. Nobody has confirmed whether Google's results page now draws properly.
  • Cookies are saved on a clean shutdown. Closing the window saves the profile. A crash, a kill, or macOS Cmd+Q that skips the window close loses the new cookies of that run. After a window close, a 3-second quit watchdog ends the process. Nobody knows whether it also covers Cmd+Q.
  • Linux keys do not survive a restart with the kernel-keyring backend this build uses. See Models and settings.

What web pages can use#

  • Page storage works. localStorage, sessionStorage, cookies and IndexedDB are saved on your computer and are still there after a restart. IndexedDB can read through an index, step a cursor and refuse a second record with the same value in a unique index. Servo could not do those, so Ferrite carries its own copy of some engine parts to add them. The Cache API (caches) is added by Ferrite on top of IndexedDB. A unique index made after records already exist is not checked against them.
  • Newer CSS works. :has(), :nth-child(n of S) and @scope are on in the engine. Container queries (@container and the cqw units) are added by a Ferrite script that reads the page's style sheets and switches the rules by measuring the containers. A rule applies below any container whose condition holds, not only the nearest one, and style() queries are not supported.
  • Service workers work for what the page itself asks for. A page can register a worker, wait for it to install and activate, message it, and have its own fetch() calls answered by the worker's fetch event, with the Cache API inside the worker. Page loads, images, scripts and style sheets are not sent through the worker, and there is no push and no background sync. The worker runs inside the page, so it does not keep working when the page is closed.
  • The Popover API is added by Ferrite (popover, showPopover, light dismiss, :popover-open in a page's own <style> tags).
  • SharedArrayBuffer can be shared with a worker. Atomics.wait works in workers and a shared WebAssembly.Memory can be sent. A compiled WebAssembly.Module cannot be sent to a worker yet, so programs that start threads the Emscripten way still fail.
  • Camera, microphone and screen sharing work in the media build, and you decide. A site that asks gets one card from Ferrite (not from the page) that says which site wants what. Screen sharing is asked every time. If the AI agent is working in the tab, a permission you gave that site before is not used, the card says so, and you are asked again. A bar across the page shows while anything is being captured and has a "Stop sharing" button. Settings lists what you allowed or blocked and lets you forget it. Checked on Linux with test sources and with a real screen under a virtual display; not checked with a real camera or microphone, on a Mac, or on Windows. Screen sharing on Wayland is not done.
  • Video and audio play only in the media build. just run-media builds Ferrite with GStreamer. Then a plain video file plays and paints, audio plays, and WebRTC peer connections work. The release page also lists files ending in -media, built by the same automatic run. On all three systems they carry their own GStreamer, so you do not install it. Before a Linux media package is published, the automatic run starts its video test in a clean Ubuntu with no GStreamer, and it must pass. The Windows packages also carry the Visual C++ files Windows needs, so they start on a clean Windows.
  • Streaming players work in the media build (Media Source Extensions). Pages that feed their video player through MediaSource and SourceBuffer (the way YouTube and most streaming sites do) now have what they need. Fragmented MP4 (H.264, H.265, AV1, AAC) and WebM (VP8, VP9, AV1, Opus, Vorbis) are read; seeking, running out of data, switching quality in the middle of a stream, timestampOffset and appendWindow work. hls.js, dash.js and Shaka Player each play a test stream, seek in it and reach its end. YouTube cannot be reached from the machine that builds Ferrite, so it is tried on the owner's Mac. There it loads and plays. One player error found there is fixed in this release: YouTube set the video's length a few milliseconds inside its last frame, and Ferrite refused that where other browsers accept it. Not done: encrypted media (no Widevine, so paid video that needs it will not play), live streams were not tested, and ManagedMediaSource and media sources in workers are missing. Tested on Linux only.
  • Google Meet does not join a call yet. Two causes are fixed in this release: SVG shapes now have getTotalLength(), getPointAtLength() and getBBox(), and Ferrite's service workers now send the header Google expects. Meet showed two more errors after those, and whether they still stop a call is not known until it is tried again.
  • Other newer web features are missing (WebGPU, Gamepad, the Navigation API and more); the list is in docs/TO-DO.md, rows T-312 and T-315.

Speed and the renderer#

  • Speed on real hardware has not been measured. One thing was measured, in the headless engine on Linux. It is the fix that stopped the shell from copying and re-uploading the whole page on every tick. Two things were not measured anywhere. They are the thread pools sized to the core count, and the release profile as just run-fast builds it. An on-disk HTTP cache was tried and then removed. The engine's disk cache stores only entries pushed out of memory, and it did nothing across restarts.
  • There is no GPU renderer. It was tried and removed. On an Apple M1 (a CI build) its self-test passed, but with it Google never finished loading and could not be scrolled or clicked, and the engine's WebGL thread panicked. With the CPU renderer, GitHub worked fully. Pages are drawn on the CPU everywhere, so they may be slower. FERRITE_WEBGL=off turns WebGL off if a page's WebGL ever freezes it. WebGL 2 is off by default (FERRITE_WEBGL=on turns it on) because of a known engine bug on macOS. Ferrite has two fixes for it in vendored engine crates, and neither has run on a Mac.

Logs and builds#

  • collect-logs is a bash script. Windows now writes ferrite.log too (in %LOCALAPPDATA%\Ferrite\logs), but the script that gathers it is bash. It works on macOS and Linux only. The stack sample works on macOS only. The script is bundled in the macOS app only.
  • Unsigned builds. Each platform asks you to confirm the first launch. On macOS, the app has an ad-hoc signature but is not notarized.
  • Builds are older than the source. A manual CI run makes the builds. The release page names the commit that a build came from.
  • The macOS app has crashed for the owner. It was reported to crash again and again, to need a force-quit, and to fail on some pages. Ferrite added a log file, a quit watchdog, crash reporting and the DevTools Engine tab so that this can be diagnosed. None of them has run on that Mac yet. The project has not checked how the packaged Windows and Linux builds behave on real machines.
  • A build without Servo cannot draw pages. It is for the interface, the defense and the evaluation.
  • No LICENSE file yet. The crate manifests say MIT OR Apache-2.0. That choice was carried over from an earlier file's stated intent. The project owner has not confirmed it.

What is explicitly not claimed#

  • That Ferrite makes AI browsing safe, or safe in production.
  • That any model resists prompt injection because it sits behind Ferrite.
  • That the evaluation holds for attacks, tasks or sites outside the corpus.
  • That users will reject unexpected actions. The headline numbers assume they do.
  • That the defense survives an attacker who has read it. The corpus was tuned against. Nobody has run a held-out, adversarial evaluation written by other people.
  • That the sanitizer finds injected text in general. It finds only the patterns it has.

What would make the claims stronger#

  • A corpus written by other people and held back, run once. The authors would not see the defense brief.
  • More real models as the agent, with each result reported. The live run so far covers one model. It shows how often that model takes the bait and what the defense contains.
  • A study of how real people answer consent prompts.
  • The model prediction layer measured on its own, with the sanitizer and the rules off.

If you can do any of these, or you find an attack that gets through, open an issue. A bypass is a result. The corpus is the place for it.