Benchmark library
Explore AI benchmarks.
Find an evaluation for the agent you are building. See what it tests before you open it.
6 benchmarksOpen repositories
τ-Voice
Sierra Research · Voice agents
Full-duplex voice-agent evaluation on grounded customer-service tasks across airline, retail, and telecom. Tests: Interruptions · turn-taking · policy adherence · tool use
March 2026
Mortgage Origination
Benchmax · Finance agents
Income and asset verification for mortgage agents reviewing borrower files. Tests: Income · assets · debts · source reconciliation
August 2026
Clinical Documentation
Benchmax · Healthcare agents
Evidence-grounded clinical notes generated from a complete visit transcript. Tests: Grounding · diagnosis support · medication accuracy
August 2026
Property Underwriting
Benchmax · Finance agents
Commercial property underwriting from submission intake to the appetite call. Tests: Source quality · conflicting facts · risk completeness
August 2026
Legal Research
Benchmax · Legal agents
Legal research from a partner’s request to the final memo. Tests: Authority · citation quality · issue coverage
August 2026
Accounting Close
Benchmax · Finance agents
Month-end close review before a controller signs the package. Tests: Reconciliation · cutoff · accruals · variance explanations
August 2026
No benchmarks match this search.