Home / Docs / Evaluation

Reproduce it

You should run the evaluation again, not take it on trust. The default way needs no network, no API key and no GPU. It is deterministic. The same corpus gives the same numbers, except for the wall-clock timing fields.

Some words on this page need a short explanation. The full list is on Words we use.

Fingerprint
The list of tools and websites a task is expected to need.
Guard
The check that stops an action outside that list.
Dry run
A practice run that does nothing real.

What you need#

  • A recent Rust toolchain. You can also use just. Every recipe below is a one-line cargo or python3 command that you can run directly.
  • The repository: git clone https://github.com/rayanjainn/Ferrite-Browser.
  • Nothing else. The evaluation crate has no Servo in it. So none of this builds the browser engine.

Run the evaluation#

$ just eval
# or:
$ cargo run -p ferrite-eval --example eval

This runs every case under all four modes (the 938-case corpus, 3,488 runs). It appends one entry per run to a hash-chained audit log. It verifies the chain. It writes the report to target/eval-report/ (or $FERRITE_EVAL_OUT_DIR):

EVAL_REPORT.md
The metrics tables: per mode, per tier, mode-pair tests with Holm correction and effect sizes, and the consent-policy sweep.
eval_report.csv
One row per run, for your own analysis.
audit.db
The audit log from the run. The run verifies its chain.
corpus.db
The SQLite store of the cases and their run records.

Unset your model variables first

The repeatable numbers come from the rules-only fingerprint, with no model. Suppose FERRITE_MODEL_SMALL, FERRITE_MODEL_MAIN and a provider key are set in your environment. Then the fingerprint's prediction layer makes live calls instead. The run is then no longer repeatable. The scripted agent never calls a model either way.

Run the runtime-guard experiment#

$ just guard-eval
# or:
$ cargo run -p ferrite-eval --example guard_eval

This runs the compromised real run against a dry run that cannot see the attack. The guard checks each action. It writes target/eval-report/GUARD_REPORT.md. The report shows the attacks that succeed without and with the guard, for each stratum, and the benign tasks the guard blocked.

Run the test suite#

$ just test
# or:
$ cargo test --workspace

This includes the component red-team suites (sanitizer, scopes, rule layer, guard, dry run, audit-log tamper matrix) and the corpus validation. It needs no network and no key. A test that would touch a live provider is marked #[ignore]. It runs only on request (just test-live).

Inspect a single case#

$ cargo run -p ferrite-eval --example inspect_case -- \
    crates/ferrite-eval/tests/corpus/c11_offscreen_scope_escalation.json

This runs one case file under every mode it defines. It prints each layer's caught or missed verdict, and the final outcome. It also says whether that matches what the case's authored ground truth implies. It is the quickest way to see why an attack was or was not contained.

Regenerate the corpus#

$ python3 scripts/gen_redteam_corpus.py            # or: just redteam-corpus
$ python3 scripts/gen_redteam_corpus.py --check    # or: just redteam-corpus-check

This writes the 909 generated cases to crates/ferrite-eval/tests/corpus_redteam/. It uses the matrix described in Evaluation. Labels come from the capability lowering table, not from running the defense. corpus_redteam_validate checks that table, and every task's capability set, against the real code. --check fails if the files on disk differ from a fresh generation. On the current tree, it reports that all 909 files match.

Check the AgentDojo import#

The 1,046-case AgentDojo corpus (see Evaluation) is checked in. The test cargo test -p ferrite-eval --test agentdojo_full_validate verifies its structure offline. To verify that it matches AgentDojo, you need the dataset. It is not in this repository:

$ git clone https://github.com/ethz-spylab/agentdojo /tmp/agentdojo
$ git -C /tmp/agentdojo checkout 089ed468cf3ed0322acc66b0211f26d9d90dbf60
$ python3 scripts/import_agentdojo.py --src /tmp/agentdojo --check

The script parses AgentDojo's task files. It never runs any of AgentDojo's code. Without --check, it rewrites the corpus. An importer run against another commit refuses unless you pass --allow-other-commit. You must not commit that output.

Try a defense mode in the app#

The app reads the defense mode once at startup, from FERRITE_DEFENSE. The values are on (the default), off, sanitizer_only or loop_only. These are the same four conditions that the evaluation compares. So you can see the difference on a page you control.

The optional live-provider runs#

There are two ways to use a real model. Neither is repeatable byte for byte. A live model's answer is not fixed the way a scripted corpus is. Neither is ever mixed into the headline numbers.

The fingerprint's prediction layer, in the same evaluation:

$ FERRITE_MODEL_SMALL=<a tag your account serves> \
  FERRITE_MODEL_MAIN=<a tag your account serves> \
  OLLAMA_API_KEY=<your key> \
  cargo run -p ferrite-eval --example eval

just models lists the tags your key can reach.

A real model as the agent is a separate runner. It has batches, a cost preview and resume. See Live evaluation. The owner has run it once, on gemma4:31b. The results are in Evaluation: Live results. The record gives these settings for that run:

  • Provider ollama.
  • Model gemma4:31b for both roles.
  • Corpus all (1,984 cases), seed 1.
  • Modes off and guard.
  • Several batches of 100 cases, with a 2000 ms pause.

The record does not give the call cap. The command below is built from those settings. It is not copied from the owner's shell:

$ cargo run --release -p ferrite-eval --example live_eval -- \
    --provider ollama --model gemma4:31b --corpus all --seed 1 \
    --modes off,guard --batch-size 100 --max-calls <N> --pause-ms 2000
$ cargo run --release -p ferrite-eval --example live_eval -- --report

Use --base-url if you use a local Ollama server. Your numbers may differ from the record, because a live model's answer is not fixed.

What “reproduced” should mean

Reproducing the default run confirms the computation. The stated corpus, run through the stated code, gives the stated numbers. It does not confirm that the corpus is typical of real attacks. For that, the next step is a held-out corpus. Someone who has not read the defense should write it.