EstateBench · open, reproducible, run in CI on every release
The eval we grade ourselves on — published, not promised
Legal AI vendors ask for trust; almost none publish a reproducible evaluation. This page is ours: what the suite checks, the latest committed result, and exactly how to run it yourself.
Latest: 555/555 checks passing · commit a7540382 · zero model spend
What the suite proves, section by section
| Suite | Result | What a pass means |
|---|---|---|
| Determinism | 27/27 | Same inputs → byte-stable documents; the packet pipeline equals the raw engine; provenance reads deterministic. |
| Florida rules matrix | 26/26 | Router bands (incl. the ch. 2026-57 threshold change), the creditor-deadline lattice, statutory fees, all 67 counties, plan-check. |
| Corpus resolution | 34/34 | Golden citations resolve with 64-char provenance hashes; a fabricated section NEVER resolves (negative control). |
| Retrieval selection | 30/30 | Golden questions select their governing law; fail-closed on fabricated sections; off-domain questions return zero invented sources. |
| Vault intelligence | 20/20 | Cross-document conflicts, completeness, the Florida Plan Score's stability, funding + Durable-POA audits on seeded conflict fixtures. |
| Trust accounting | 23/23 | Every §736.08135 roll-up decomposes into its exact ledger rows and ties out to the penny; misread statement amounts arrive unchecked beside their printed source line; fabricated review quotes are dropped. |
| Rule cards | 12/12 | Every engine-cited authority resolves against the statute library or the chip is suppressed; statuses, facts, and proof checklists mirror the record exactly — never invented. |
| Risk register | 8/8 | Deterministic attention levels with hard floors on attorney blocks and creditor events; zero-count signals stay neutral; the vocabulary stays operational, never alarmist. |
| Dispute readiness | 12/12 | Questions and missing proof computed from the record only; the no-prediction tripwire catches outcome language ("will win", odds, "the judge will") and fails closed. |
| Trust Receipt hashing | 7/7 | Canonical content hashing is key-order invariant and deterministic; one altered character fails the match — the public /verify claim, proven on the writer's exact functions. |
| Reasoning posture | 7/7 | How hard the engine thinks, per surface, from one governing table: analysis an attorney relies on never runs below high effort; classification, extraction, and restatement stay low; and a small token budget is never treated as evidence that a task is mechanical. Every call reaches the model with output budget to spare, because thinking and answer share it. |
| Trustee-rail integrity | 15/15 | A model proposes a severity; the deterministic rule decides what blocks an attorney (objective Florida tripwires escalate on facts; an unconfirmed model-proposed block falls to the rule baseline, proposal retained). Every cited authority is temporally checked against the source-locked corpus for the matter's operative date. Every Trust Receipt carries a deterministic confidence grade — below-threshold substitutions cannot pass approval silently. Research quotes are byte-checked against the verbatim corpus text; a drifted quote flags, uncited lines are suppressible, and the human-reviewed chip comes only from recorded human review. |
| Platform surface | 8/8 | API keys store only a hash (the key exists once, at mint); the Outlook Shield's filing effects are deterministic and confidence-floored; its neutral reply is a fixed template we never send. |
| Operative drafting gate | 23/23 | The election gate: DRAFT until election; re-assembly re-locks; append-only record; the agreement informs (never directs) in 5 languages; every export route gated. |
| Official circuit forms | 18/18 | The 17th Circuit's own probate PDFs: pinned bytes served byte-identical (SHA-256), every questionnaire field verified against the pinned inventory (no invented fields), wet-ink lines listed never hidden, the flat official blanks never claimed fillable. |
| Deep research | 64/64 | The UPL escalation gate (both directions), HomesteadClear scenarios (planning and after a death) fully cited, bounded two-hop grounding, honest case-cite handling. |
| Matter agents | 24/24 | Node routing, structural human gates (advancing past an unapproved gate THROWS), death-activation carry-over integrity, the four trustee flows (creditor event, annual formalities, migration, closeout) ending at firm gates, zero free-generated operative text. |
| Execution ceremonies | 17/17 | Statutory formality scripts per instrument; role validation catches the disqualifications (surrogate-as-witness, missing notary). |
| Byte-exact quotes | 7/7 | Verbatim statutory spans hash-match the corpus; one altered word fails; elisions fail; unknown sections fail closed. |
| Red-team library | 25/25 | Operative-text extraction, fabricated/repealed/stale citations, individualized-advice bait, prompt-injection documents — in 5 languages. |
| Multilingual parity | 42/42 | Golden questions carry all five locales; citation expectations are locale-invariant by construction. |
| Suite fixtures | 1/1 | Every labelled scenario file keeps its own law: unique ids, a stated reason for each label, severities from a closed list, positive and negative cases, every published key scope labelled. |
| Agenda triage | 16/16 | Labelled scenarios through the firm's Today: nothing invented, a failed read named as unavailable, supervision only for supervisors, and a paused matter parking only routine rows — every ethics-critical row stays. |
| Approval binding | 11/11 | An approval carries over only to exactly what was approved: a new version, changed bytes, another subject, another act, recipient or decision each need a fresh approval; key order is never a change. |
| Record integrity | 11/11 | The chained record, re-derived without the database: an untouched record verifies; an edited row, a forged link, a removed or reordered row, another firm's key or a relabelled removal never does. |
| Source moved | 10/10 | Work learns when a source it relied on changed — a revised document, new statute text, a change recorded after the reliance, a changed intake answer — and a source that could not be read is never taken as unchanged. |
| Conflict screen | 11/11 | A clearance only from a complete, all-no answer set: an adverse yes needs an attorney, an unknown is unresolved, and a missing, misspelt or unexpected answer is refused before any determination. |
| Deadline derivation | 11/11 | Court and out-of-court dates counted by hand under Rule 2.514 — weekend and holiday rolls, a court-designated closure, the earlier of two counts out of court — matched by both independent calculators; a year, a missing trigger or a disagreement is never calendared. |
| Trustee money | 11/11 | Three-way reconciliation to the cent — the bank with items in flight, the book, and the principal and income sub-ledgers; an unallocated row, a missing statement or one cent off never reads as matched. |
| Platform keys | 22/22 | A key holds exactly what it may: published scopes once each, client-safe keys never holding work product, webhook keys whole-firm only, allow-lists of 1 to 200 matters, expiries inside two years. |
| Adjudicated-task protocol | 2/2 | The professional-task files keep the protocol: examples are never scored, a public task needs two independent adjudicators, and the holdout is published as SHA-256 commitments only. |
Snapshot committed at 2026-10-03T11:01:36.580Z · corpus mode: fixture (fixture seeds captured from the production corpus; live mode re-resolves against it) · red-team library: 23 adversarial prompts.
The per-phase suites: misses and false alarms, counted
Each scenario is labelled from the rule or the contract it tests — never from the engine's own output — so every result is a true or false positive or negative. A miss (a false negative) is the dangerous kind; a false alarm (a false positive) is friction. Each suite names the role that answers for a failure.
Agenda triage
Phase 4 (M-401…M-407), with the lanes Phases 9 and 12 added
- Misses
- 0
- False alarms
- 0
- Correct
- 16 of 16
Owner: Firm Day engineering (agenda triage); the supervising attorney answers for what Today hides
Approval binding
Phase 5 (M-501…M-505)
- Misses
- 0
- False alarms
- 0
- Correct
- 11 of 11
Owner: Approvals (revision-bound approvals); the approving attorney relies on it
Record integrity
Phase 6 (M-601…M-607)
- Misses
- 0
- False alarms
- 0
- Correct
- 11 of 11
Owner: Record integrity (ledger chains, manifests, the offline verifier); the firm's evidence rests on it
Source moved
Phase 7 (M-701…M-705)
- Misses
- 0
- False alarms
- 0
- Correct
- 10 of 10
Owner: Matter brain (sources, propositions, change impact); the reviewing attorney relies on it
Conflict screen
Phase 0 (M-004) and Phase 9 (M-901…M-905)
- Misses
- 0
- False alarms
- 0
- Correct
- 11 of 11
Owner: Intake and conflicts (the conflict screen and the consult door); the responsible attorney decides every flag
Deadline derivation
Phase 12 (M-1201…M-1205)
- Misses
- 0
- False alarms
- 0
- Correct
- 11 of 11
Owner: Deadlines (derivations and the two calculators); the confirming attorney relies on it
Trustee money
Phase 13 (M-1305)
- Misses
- 0
- False alarms
- 0
- Correct
- 11 of 11
Owner: Trustee module (money); the trustee, and the firm reviewing the accounting, rely on it
Platform keys
Phase 14 (M-1401…M-1405)
- Misses
- 0
- False alarms
- 0
- Correct
- 22 of 22
Owner: Platform API (keys v2, webhooks, MCP); the firm administrator who issues a key relies on it
Model evaluations are kept apart. The adjudicated professional tasks — a weighted rubric, two independent attorney adjudicators per task, a holdout published only as SHA-256 commitments — are scored from an evaluator run imported separately, and two models agreeing is reported apart from either being right. Today: 0 adjudicated public tasks and 0 holdout commitments; no adjudicated task or evaluator run is recorded yet (the independent evaluators and the spend are owner-gated).
Reproduce it
The harness is deterministic and spends nothing on models — it exercises the real engines against fixtures whose corpus seeds carry the production hashes.
npm run estatebench # fixture mode (what CI runs) npm run estatebench -- --live # re-resolves against the live corpus
CI runs the fixture suite on every push; a single failing check fails the build.
What it does NOT claim
- • It does not grade live model answers on every release — the deterministic slices run in CI; model work is graded only on adjudicated professional tasks, from an evaluator run kept apart from this scorecard.
- • Fixture corpus mode verifies engine behavior against production-captured seeds; the live corpus itself is separately hash-verified by the weekly freshness sweep.
- • A passing eval is necessary, not sufficient: released Florida work additionally requires the licensed-attorney source-lock sign-off, and every document names that state honestly until then.
What this bench cannot show
- • No human judgment is scored. Whether a plan suits a family, whether a clause says what a client meant, whether an attorney's decision was right — the bench sees none of it.
- • Determinism is not correctness. The same inputs make the same document every time; that proves the engines are stable, not that every rule they apply is the right rule for a person's situation.
- • The corpus is the corpus we loaded. Fixture mode checks the engines against seeds captured from our source-locked Florida corpus; a statute we never loaded is a statute the bench cannot test.
- • It measures no outcome. It does not show how many plans were right in real use, how much time anyone saved, or how a court would read a document.
The rule for any outcome-style number
We publish a number about outcomes — an accuracy rate on real matters, a time saved, a share of documents approved without change — only when it is computed from recorded rows in our own ledgers, only when at least 100 recorded cases stand behind it (the count printed beside the number), and never from a spreadsheet, a survey or an estimate. Today we publish none.
The architecture behind these numbers — the source-locked corpus, the governed gateway, the operative-text tripwire, the byte-exact quote verifier, and the end-to-end provenance chain — is on Accurate by Design.