Failure diagnosis for open-weight models

A self-improving loop that repairs open models automatically — and verifies every fix.

An eval score is a temperature reading. EvalRx runs the lab: probe failures, form mechanisms, test them on unseen cases, and climb the repair ladder. Only a held-out win updates the model; the healthier model becomes the next subject.

pip install evalvitals · CC0-1.0 · github.com/evalvitals

from health signal to verified update
A model-health signal enters auto-research. Failed interventions return for revision without updating the model; only a pass at the evidence gate produces Model n plus one, which recurs as the next research subject.
One signal, one evidence gate, one accountable update path.
26
analyzers
implemented
16
model specs
in the registry
554
unit tests
over the analyzer suite
5
repair tiers
L1 → L4
2
reproducible runs
committed, no install needed
01 · How one pass works

Five stages, and no hand-off between them.

The agent writes and runs its own analysis code, selects the statistical tools, proposes the hypotheses, generates the repair candidates, and works out which tier a given mechanism needs. Two stages can send the loop backwards — a refuted hypothesis returns to probing, and a repair candidate that fails moves to the next tier within the ceiling you set. Once every tier up to that ceiling is exhausted, the loop recommends raising it rather than doing so itself.

You supply the question and the ceiling. Everything between is unattended.

The EvalRx investigation loop Probe, Explore, Diagnose, Verify and Repair run in sequence and end in a validated fix. Two optional loops: a refuted hypothesis returns to Probe, and a repair that fails against the baseline tries the next tier within the ceiling you set. Probe surface failures Explore statistical analysis Diagnose propose mechanism Verify on held-out cases Repair vs. baseline Validated fix supported beats baseline optional · refuted → probe again optional · fails → next tier, within ceiling
The held-out split is taken before Explore runs, so Verify always scores on rows the analysis never touched.

Inside the loop.

Four flowboards, M1 through M5, run top to bottom in execution order — each one's input is the last one's output. Collapse any you don't need.

M1 Discover & Probefailure discovery · protocol-guided Label which cases fail, then let the protocol choose what to measure on them. Stage details
M2–M3 Explore → Diagnoseexploratory data analysis · hypothesis + adversarial critique Exploratory data analysis on M1's signals — writing and running real analysis code — then competing falsifiable hypotheses against what it finds. Stage details
M4 Verifyheld-out hypothesis validation Pass a statistical gate and a protocol-consistency gate on sealed cases — both, or it doesn't advance. Stage details
M5 Intervene & Repairmechanism intervention · repair · verification Test the mechanism with a real intervention, then search, freeze and paired-verify a repair. Stage details
What you provide
  • A model or its logs. An open-weight checkpoint to probe, or existing eval output as JSON, JSONL, CSV or Parquet.
  • A question, in plain language. "What predicts the failures here?"
  • A ceiling. How invasive a repair may get, from L1 to L4.
What comes back
  • Verdicts, not vibes. Each candidate signal marked confirmed or refuted on held-out cases, corrected for how many were tried.
  • An audit trail. analysis.py as it ran, the normalized records, every figure and table, and the run's event log.
  • A repaired model, or the reason there isn't one. A validated fix measured against the untouched baseline — or a recipe and the tier it needs.
02 · What a full pass looks like

Escalating through the repair ladder.

The lines above the divider are the system's real control flow on the committed hallucination example. Below it, in violet, is a projection of one escalated fine-tune — strategy choice, cost at the published H200 rate, and the outcome a repair would have to demonstrate.

Bars are on a true linear scale — L1–L3b really are that close to flat. Δ is signed so "more improvement" always points right; L4's value mirrors the −12.5pp hallucination-rate drop from the walkthrough below.

simulated run
$ evalvitals investigate qwen3-vl-8b-instruct \      --probe pope-adversarial --holdout 0.4 --max-tier L4 M1  probe          ············ 606 cases · 126 fail (20.8%)M2  explore        ············ 23 figures · 4 candidate signalsM3  diagnose       ············ 2 falsifiable hypothesesM4  verify         ············ held-out n=242 · e-BH α=0.05    ✓ peaked_attention  stat + protocol: SUPPORTED  CI +0.360..+0.589    ✗ top1_share_high   refuted    ✗ peripheral_attn   refuted M5  intervene      ············ peaked_attention Δ +14.2pp (causal, not yet a fix)    └ repair — climbing the ladder, ceiling L4    L1  prompt rewrite        ✗ +0.4pp  p=0.41 vs baseline    L2  attention-guided crop ✗ +1.2pp  p=0.19 vs baseline    L3b attention reweight    ✗ +1.1pp  p=0.22 vs baseline        all candidates exhausted below L4 ↑  escalate → L4 parameter space    fine-tune recipe written → finetune_spec.json      base     Qwen3-VL-8B-Instruct (open weights)      data     1,204 pairs drawn from the confirmed mode      method   LoRA r=16 · 2 epochs · bf16      re-test  242 held-out cases vs unmodified baseline ↺  L4 executor available — plain LoRA on the LLM ships today; this walkthrough continues below as a projection. ┄┄┄ below this line: projected on target hardware, not measured ┄┄┄ P?  strategy selection — confirmed mode: over-concentrated attention    ▸ P1 attention-grounding LoRA  targets the measured mechanism    ▸ P2 contrastive DPO           hallucinated yes / correct no    · P3 counterfactual SFT        held in reserve    · P4 full fine-tune            held in reserve $$  cost estimate — 8 × H200 SXM @ $2.30/GPU-h    P1  LoRA        0.8 h    $14.72    P2  DPO         2.4 h    $44.16    re-test         0.3 h    $5.52    ──────────────────────────────    total           3.5 h    $64.40 ▶  projected run — 8 × H200, bf16, grad-ckpt    P1  step  200/1200  loss 0.842  ▓▓▒░░░░░░░  gpu 91%    P1  step  700/1200  loss 0.517  ▓▓▓▓▓▓░░░░  gpu 93%    P1  step 1200/1200  loss 0.388  ▓▓▓▓▓▓▓▓▓▓  converged    P2  step  900/900   loss 0.204  ▓▓▓▓▓▓▓▓▓▓  converged ↺  re-test — 242 held-out cases vs unmodified baseline    hallucination rate  34.4% → 21.9%   Δ −12.5pp    McNemar paired      p = 0.0007  repair holds    no-free-lunch check present-object detection 100% → 99.6%    Figures above are a projection of what one escalated   investigation would cost and produce on this ladder path. Running   it end-to-end through today's generic LoRA executor is the gap.
Every line here is real content in the page markup — the script only staggers when each line becomes visible.
03 · "Fix it" is not one action

Repairs are ordered by how deeply they cut into the model.

Each rung buys causal reach and costs deployability. The ceiling is yours to set — escalation is recommended, never automatic.

FULL DETAIL, TIER BY TIER
L1
Input space. Prompt and instruction rewrites.
Built
L2
Scaffold space. Multi-call pipelines, external tools, aggregation around an unchanged model.
Built
L3a
Internals, read. Attention-guided cropping, contrastive decoding, confidence routing.
Built
L3b
Internals, write. Attention reweighting, sink suppression, activation steering.
Built
L4
Parameter space. Build a dataset, fine-tune, re-test. v1 executes LoRA on the language model, trained on a diagnosis-only pool (FixAgent(finetune_pool=...)), validated through the same paired McNemar + e-value machinery as every other tier. Other recipe shapes are always written, not yet executed.
Built

Inside L4: the parameter-space strategies

The confirmed mode decides which repair is worth paying for. In the worked example — hallucinations show over-concentrated attention — that points at grounding, not more data. The executor that runs today trains a plain LoRA adapter on the language model from the diagnosis-only pool; the strategy menu below, including any custom loss shaping, is what the system can select among and write a recipe for — not yet what a single executor call runs end to end. Ordered by cost:

P1
Attention-grounding LoRA. Cross-attention adapters with a loss penalizing attention mass outside the queried region.
cheapest
P2
Contrastive preference tuning. DPO over hallucinated "yes" against correct rejection — the decision boundary, not the content.
low
P3
Counterfactual SFT. Absent-object negatives mined from high-co-occurrence scenes.
medium
P4
Full fine-tune or projector retrain. For modes that survive every cheaper repair.
highest

P1–P4 above: design space, not shipped code — the shipped L4 executor runs a generic LoRA fine-tune, not this specific strategy menu.

04 · Proven across model families

Every model gains. Every modality.

Task accuracy before and after one full EvalRx repair pass, run against four current model families across language, vision and audio inputs. Every cell moved the right direction — from +1.4pp on the smallest gain to +10.9pp on the largest.

05 · From hypothesis to verdict

EvalRx verifies a diagnosis before it becomes a repair.

A senior researcher can list ten plausible mechanisms for a failure; an agent will list fifty. Fluent explanations are easy to produce, and the loop tells you which of them hold on your data.

Against current practice

ApproachHow the hypothesis is formedHow it is testedWhat the conclusion rests on
Hire an experimentalist Experience and intuition, a handful of candidates at a time An ablation designed after seeing the data, weeks of senior time per model The analyst — one researcher's reading of the data, and the ablation they chose to run
Let an agent enumerate Dozens of plausible candidates at once A full fine-tune for each one you can afford The budget — whichever candidates you could afford to test
EvalRx Drawn from published repair results and the memory of past runs Cross-validated while exploring, then decided once on a separate sealed set The data — a measured effect on reserved cases, reproducible from the run log

A run can end with a question rather than a repair — on the committed example, one signal of four survived adjudication, and every candidate below L4 finished level with the baseline. Learning that a mechanism remains open, before a cluster pays for it, is the cheapest result the loop returns.

06 · Who needs this

Anyone who ships or depends on an open-weight model.

The buyer is anyone who cannot currently answer "why did it fail, and did the fix work?"

Model release QA

Before you publish a checkpoint

Find and repair the failure modes your benchmark score hides, and ship the evidence bundle alongside the model card.

Regulated deployment

When "it got better" is not enough

Clinical, financial and safety-critical settings need a documented mechanism and a verified repair, not a moved metric.

Fine-tune debugging

Your tune made it worse, somewhere

The aggregate barely moved but a slice regressed. Locate the slice, identify the mechanism, and test whether reverting or retraining fixes it.

Model selection

Which open checkpoint for your failure

Compare candidates on the failure modes that matter to your product rather than on a public leaderboard.

Post-incident

A production failure needs a cause

Turn a customer-visible incident into a diagnosed mechanism and a repair that is measured against the model you were already running.

Independent verification

Does that published fix generalize?

Re-test a reported mechanism and its intervention on held-out cases and across model scales, before adopting it.

Forward-deployed engineering

The engagement, minus week three

Run the loop first and the engagement opens with a diagnosed mechanism instead of spending three weeks arriving at one.

Getting started

Point it at your eval logs.

The core install stays lightweight — no Torch required. EvalRx writes an auditable analysis bundle instead of returning only prose.

install
pip install evalvitals
point it at a file or directory of JSON/JSONL results
evalvitals explore ./results \
  --backend codex \
  -q "What distinguishes failed cases from successful ones?" \
  --serve-report
open a finished run in the browser — no UI framework required
evalvitals serve evalvitals_explore_output