recordroom

Discover, collaborate, and build your own benchmarks.

Written standards for AI agents doing real work — one-sentence evals, case corpora, judged runs with cited evidence. All files, all forkable.

benchmax/revenue-coachingv0.1.0

End-to-end evaluation of agents that analyze sales calls, coach reps, and update CRM state.

5 tasks · 1 environment updated 2026-08-19 · CC-BY-4.0
recordroom/property-underwritingv0.2.0

What a commercial property underwriting agent must get right, from submission intake to the appetite call.

no published run yet 21 evals · 29 cases · 50 recorded outputs synthesized + UNDERWRITE sessions · CC-BY-4.0
recordroom/clinical-documentationv0.2.0

What a clinical documentation agent must get right turning a visit transcript into the note a clinician signs.

no published run yet 14 evals · 14 cases · 14 recorded outputs synthesized corpus · CC-BY-4.0
recordroom/legal-researchv0.1.0

What a legal research agent must get right, from the partner's request to the memo it returns.

no published run yet 14 evals · 14 cases · 14 recorded outputs synthesized corpus · CC-BY-4.0
recordroom/accounting-closev0.1.1

What an AI close agent must get right before a controller signs the month-end close package.

no published run yet 14 evals · 14 cases · 14 recorded outputs synthesized corpus · CC-BY-4.0
benchmax/mortgage-originationv0.3.0

A synthetic document-review benchmark for testing whether mortgage AI systems classify, reconcile, and explain borrower-file evidence consistently.

14 tasks updated 2026-08-14 · CC-BY-4.0
benchmax/tau-voicev1.0.0

Full-duplex voice-agent evaluation on 278 grounded customer-service tasks across airline, retail, and telecom.

278 tasks updated 2026-08-11 · Apache-2.0
Collaborate

Every spec is markdown, every run is JSON, and every verdict cites its spans. Read a sentence you'd write differently? That argument is the product working.

Disagree with a sentence →
Build your own

The format is one document: FORMAT.md. Start from the template — or tell us the workflow and we'll draft the spec with you.

Start a benchmark →