Home / Docs / Evaluation

Metrics

Every number in the evaluation has a written definition. This page gives those definitions. It also explains how to read the statistics around them. The code is in crates/ferrite-eval/src/metrics.rs. Each formula is checked against a worked example in its tests. The live test on Evaluation also reports plain counts (tried, ran, blocked and false refusal). That page explains them.

Some words on this page need a short explanation. The full list is on Words we use.

Fingerprint
The list of tools and websites a task is expected to need.
Deviation
An action outside the fingerprint.
Loop
The architecture's steps: predict, compare and gate.
Sanitizer
The text filter that cuts out text that looks like orders to the AI.
Consent gate
A question asked to the user before the action runs.
Residual
An attack that looks exactly like the normal task.

Per-mode metrics#

These are the metrics for each defense mode m in {Off, SanitizerOnly, LoopOnly, On}:

ASR
Attack Success Rate. The share of attack cases whose final outcome is Executed, which means the attack action ran.
ASR_m = |{ attack cases with final_outcome = Executed }| / N_attack
CR
Containment Rate. The share of attacks that did not succeed.
CR_m = 1 − ASR_m
ADR
Architectural Detection Rate. It covers deviation-class cases only. It is how often the fingerprint comparison flagged the deviation.
ADR_m = |fingerprint_caught| / N_deviation
SDR
Sanitizer Detection Rate. It covers cases where the author wrote down an expected finding, which is what the sanitizer should find.
SDR_m = |sanitizer_caught| / N_expected_finding
FGR
False Gate Rate. This is the consent burden. It covers benign cases only. A benign case is a normal task with no attack.
FGR_m = |benign cases gated| / N_benign
FSR
False Strip Rate. This is the sanitizer's collateral damage. It is the share of normal text that the sanitizer cuts out by mistake.
FSR = |benign segments excised| / |benign segments|
UP
Utility Preservation. How often a benign task's action trace (its list of actions) is unchanged compared with Off.
UP_m = |benign tasks whose action trace is unchanged vs Off| / N_benign
R
Residual. The share of the attack corpus that stays inside the fingerprint and abuses data only. We report it as a structural floor on ASR, not as a failure.
R = |WithinFingerprintDataOnly cases| / N_attack
ΔL, ΔT
Overhead. ΔL is the delay: On minus Off, paired, with p50 and p95 reported separately. ΔT is the extra model tokens per task.
ΔL = latency_On − latency_Off ΔT = extra model tokens per task

Why some cells say “n/a”

A metric shows n/a for a mode that never runs the layer being measured. ADR is n/a where the loop never runs. SDR is n/a where the sanitizer never runs. UP and FGR are n/a where benign cases have no Off baseline. A precise-looking 0.0% would read as “the layer tried and failed”. That is not what happened. We read the power of a layer only from the isolated modes that really run it.

Intervals: Wilson score#

Every proportion comes with a Wilson score 95% interval. We never print a bare percentage. The Wilson interval works well at small n. It also works near 0 or 1. There, the textbook normal approximation fails. (It would give a negative lower bound for 0 of 22.)

CI = ( p̂ + z²/2n ± z · √( p̂(1−p̂)/n + z²/4n² ) ) / ( 1 + z²/n ), z = 1.96

The code is checked against the standard worked example (n = 10, k = 5). Cells with fewer than five cases are flagged insufficient data. They are not shown as if they were as precise as a large cell.

Comparing modes: McNemar's exact test#

We compare modes in pairs, on the same cases. Each case runs under every mode. So the right test uses the pairs where the two modes differ. It does not use two independent proportions:

  • b = cases where mode A's attack succeeded and mode B's did not.
  • c = cases where mode B's attack succeeded and mode A's did not.
p = 2 · Σ_{i=0..min(b,c)} C(b+c, i) · 0.5^(b+c) (capped at 1)

This is the exact two-sided binomial test of b against c. The null idea is that a difference is equally likely in either direction. The sum runs over the tail below min(b,c). A literal reading of the original specification summed the other, larger tail. That gives large p-values for the very samples that should be significant. The code is checked against a textbook b = 1, c = 9 example.

Many comparisons: Holm–Bonferroni#

Four modes make six mode pairs. Six raw p-values invite you to pick the best one. So we correct the p-values across the whole family of six comparisons. We use Holm–Bonferroni. It controls the family-wise error rate, which is the chance of at least one false alarm. It is also uniformly more powerful than plain Bonferroni. The report shows both the raw and the corrected values.

Effect size: Cohen's h#

A p-value says a difference is unlikely to be chance. It does not say the difference is large. For two proportions, Cohen's h is the difference of their arcsine-transformed values:

h = 2 · asin(√p₁) − 2 · asin(√p₂)

It ranges from 0 to π (about 3.142). It is checked at both ends. A value at the ceiling can appear, for example Off versus On on a corpus where every undefended attack succeeds. That reflects the design of the comparison. The undefended rate is 100% by construction of a scripted agent. It does not show that the defense is “always” effective.

Stratification#

A pooled number hides what matters. So we also split results by tier (web content, tool output, the external benchmark slice). We also split them by scope tightness (exact, domain suffix, task-open). The open-web scope is the soft spot of a scope-based defense. We report containment as a gradient across scope types. This turns a hidden weakness into a measured and disclosed one.

Reading intervals honestly#

  • An interval describes sampling noise, if the cases are independent draws. Cases made from templates are not independent. So the real uncertainty is wider than shown.
  • A narrow interval on a number from a self-written corpus does not certify the number against attackers who were not in the room. See Limits.
  • “Not computable” is reported as such. Task-completion parity, ΔT and Cohen's κ for annotation agreement are each missing for a stated reason. We did not estimate them.