Live evaluation
The headline evaluation is scripted. Its agent does what each case's attack says. The live runner is the other half. It puts a real model in the agent's seat. Then the question is what a model really does under an injection, and what Ferrite does about it. An injection is hidden text that tries to give the AI new orders. You run the runner on your machine, with your own keys, in batches your quota can survive.
It has been run once
The owner ran it once on a real model. The model was gemma4:31b on Ollama, in both roles, over all 1,984 cases. That gave 3,968 runs, with 0 failures. The results are in Evaluation: Live results. Before that run, every behaviour of the runner was tested offline. The tests used a deterministic mock provider and a loopback fake server. The step and token estimates that the cost preview prints are still estimates. This site has not compared them with the real run. Nothing here changes the headline numbers on the Evaluation page.
Some words on this page need a short explanation. The full list is on Words we use.
- Fingerprint
- The list of tools and websites a task is expected to need.
- Guard
- The check that stops an action outside that list.
- Dry run
- A practice run that does nothing real.
- Consent gate
- A question asked to the user before the action runs.
- Residual
- An attack that looks exactly like the normal task.
What it runs#
For each case, the runner drives the app's own agent loop. It uses the same action schema, system prompt and parser. It runs against the dry-run engine. The injection is planted in the page or tool output that the agent reads. Two roles can each be a model or not. You choose each role on its own:
| Role | llm (default) | Other |
|---|---|---|
Predictor (--predictor) | The small-tier model proposes the fingerprint's may_use set, as in the app | rules: only the rule layer |
Agent (--agent) | The main-tier model chooses actions | scripted: the worst-case script, with no calls |
The defense differs for each case. The modes in --modes set it (default off,guard):
- off
- No defense. This is the baseline.
- guard
- The runtime guard refuses every action outside the predicted fingerprint. The sanitizer is off.
- full
- The guard, plus the sanitizer on what the agent reads.
- dryrun
- Stage one only: the plan on a clean synthetic page, compared with the prediction. This shows the consent burden, which is how many questions the user would be asked.
The simulated user refuses every deviation. This is the best case for containment.
Providers and keys#
The runner reuses the app's model layer and adds no HTTP stack. So it supports what that layer supports:
--provider | What it is | Key |
|---|---|---|
gemini | Google Gemini | FERRITE_GEMINI_API_KEY, or the OS keyring (service ferrite) |
ollama | Ollama Cloud, or a local server with --base-url http://localhost:11434 | OLLAMA_API_KEY, or the keyring. A local server needs none |
mock | A deterministic stand-in with no network. It tests the pipeline and says nothing about any model | none |
OpenAI-compatible and Anthropic endpoints are not supported, because the model layer has no backend for them. Adding one means changing the model layer, not the runner. (You can reach a hosted server that also speaks Ollama's /api/chat with --base-url.)
The runner reads keys from the environment or the keyring (the OS credential store), and nowhere else. It never accepts a key as a flag. It never prints or writes a key. Error text kept in a result passes through a redactor. The redactor removes the key values the process holds. Model tags are configuration with no default. Use --model TAG for both roles. Or use --small-model and --main-model. Or use the FERRITE_LIVE_* and FERRITE_MODEL_* variables.
Cost first: --plan#
--plan prints what a selection would cost. It calls nothing, so it needs no key:
$ cargo run --release -p ferrite-eval --example live_eval -- \
--plan --provider gemini --model <tag> --corpus agentdojo
The output lists the cases and runs in the window. It lists how many are already stored. It lists the fingerprint predictions. It lists the agent steps (typical and worst case). It gives a rough token count (characters divided by four). It also says how many invocations the call cap implies. For the full AgentDojo import at the default two modes, it prints 2,092 runs and about 7,300 model calls typical (51,254 at most). That is roughly 74 invocations at the default cap of 100 calls. These are estimates from the corpus and the step limit. Only a real run shows what a model really takes, what the cache absorbs, and the provider's own tokenizer.
Batches, limits and resuming#
- Selection.
--corpus agentdojo|redteam|core|pilot|agentdojo-hand|all(defaultagentdojo),--suite agentdojo/banking,--only attack|benign,--seed S(a shuffle you can repeat, so a small batch samples every suite),--offsetand--limit. - Batches.
--batch-size Nruns the next N cases that still have work to do. Run the same command again for the next batch.--batch-index Ktakes slice K of the whole ordering instead. - A hard call cap.
--max-calls N(default 100) caps the requests that reach the provider, retries included. When the cap is spent, the case in flight is dropped and not recorded. It stays pending. The runner writes a partial-results ledger and exits with code 3. - Rate limits.
--pause-msspaces the calls. The runner retries a 429, a server error or a timeout, up to--max-attemptstimes (default 5). It waits longer each time, with random jitter. It honours the provider'sRetry-After, capped at 90 seconds by default, so it does not sleep through an hour-long daily quota. After 3 failed calls in a row (--max-consecutive-failures), the invocation stops with exit code 4. The provider is saying no, so come back later. - Failures are not scores. A call that never got an answer is stored as an error and left out of every rate. This covers a rate limit, an outage, a rejected key or a timeout.
--retry-failedredoes it. A model that answers badly is a result. This covers an action that cannot be parsed, an empty reply, or an unusable fingerprint. The fingerprint falls back to empty, and the case is scored as such. - Resume. Results go to an append-only file, one line per case and mode. The file is synced before the next case starts. It lives under
target/live-eval/results/. A crash, Ctrl+C or a closed laptop loses at most the case in flight. Run the same command again, and it skips what is stored. This works as long as the settings that change behaviour (models, modes, steps, predictor, agent) give the same hash. Model responses are also cached at temperature 0. So re-running an unchanged case is free, and it does not count against the cap.
| Exit code | Meaning |
|---|---|
| 0 | The batch finished cleanly. Go on to the next one. |
| 1 | An internal or I/O failure. |
| 2 | Usage or configuration: a bad flag, a missing key or tag. |
| 3 | The call cap was reached. Nothing is lost. |
| 4 | The provider kept failing. Wait, then run it again. |
| 5 | The batch finished but some cases failed. Run it again with --retry-failed after a pause. |
What to run first#
# 1. nothing is called; no key needed
$ cargo run --release -p ferrite-eval --example live_eval -- --plan \
--provider gemini --model <tag> --corpus agentdojo --seed 1 --batch-size 10
# 2. one small real batch (key in FERRITE_GEMINI_API_KEY or the keyring)
$ cargo run --release -p ferrite-eval --example live_eval -- \
--provider gemini --model <tag> --corpus agentdojo --seed 1 \
--batch-size 10 --max-calls 100 --pause-ms 4000
# 3. read what happened, then repeat step 2 for the next batch
$ cargo run --release -p ferrite-eval --example live_eval -- --report
First do a --provider mock --model x run once. You will see the whole flow with no key. If you have just, then just live-eval … and just live-eval-plan … pass the same flags to the same program.
The report#
--report (optionally --compare gemini:<tag>,ollama:<tag>) writes REPORT.md and report.csv from the stored results. It makes one group for each provider and model pair. Every rate has a 95% Wilson range (a range that shows how uncertain the rate is). Rows with few observations say so. The report covers:
- Attack runs for each mode: tried, executed, blocked of tried, and final answer.
- Baseline against defense, case by case: McNemar's exact test and Cohen's h.
- Benign tasks: false positives and the consent burden.
- Fingerprint precision and recall.
- Breakdowns by suite, attack class, ground truth and carrier.
- Cost: calls, cache hits, rate-limit answers and tokens.
A residual attack uses the task's own primitive at its own origin. The report counts residual attacks as a floor, not a failure. Some attacks have no attacker string to recognize. The report counts them and leaves them out of attack rates. The first report is summarized on Evaluation: Live results.
The AgentDojo corpus#
The runner's default corpus is the full AgentDojo import. It has 1,046 cases: 949 attacks and 97 benign twins. They come from AgentDojo v1.2.2 (MIT) at commit 089ed468cf3ed0322acc66b0211f26d9d90dbf60. Evaluation: AgentDojo in full says how it is built, and what it does and does not mean. These are the runner's plan counts for each suite:
| Suite | User tasks × injection tasks | Attack cases | Benign twins | Cases |
|---|---|---|---|---|
| workspace | 40 × 14 | 560 | 40 | 600 |
| travel | 20 × 7 | 140 | 20 | 160 |
| banking | 16 × 9 | 144 | 16 | 160 |
| slack | 21 × 5 | 105 | 21 | 126 |
| All | 949 | 97 | 1,046 |
What this does not establish#
- The labels come from AgentDojo's ground-truth calls, mapped onto Ferrite's primitives. They do not come from AgentDojo's own checks.
- It uses one attack template (
important_instructions) and one try per case, at temperature 0. - The environment is synthetic. “Task success” is estimated by the agent giving a final answer.
- The simulated user refuses every deviation.
- Only one real model has been run so far (
gemma4:31b). It says nothing about other models. - The numbers are Ferrite's own measurement. They are not comparable to published AgentDojo utility or security scores.