Bevaya Labs
Benchmarks
How we measure our models against frontier models on real insurance documents. Bevaya's models lead every frontier model tested on accuracy, at a fraction of the cost.
See the latest resultsLoss run extractionhigher is better
Macro accuracy
Bevaya0%
Fable 50%
Opus 50%
Gemini 3.60%
GPT-50%
93.1%
Macro accuracy
+7.8 pts above next-best model
63×
Cost advantage
cheaper than Fable 5 at same task
2.4×
Faster
than Fable 5 · 20.9s vs 49.9s
Loss run extraction
346 loss run documents read into structured claims and policy tables. Our Loss Run Model against 12 frontier models across 18 configurations.
First published run. Frontier models shown with and without reasoning where the provider offers both. Pricing confirmed July 2026.
Loss run extraction results
Cost vs % Perfect
Log scale · per sample (USD)
Latency vs % Perfect
Log scale · wall-clock seconds per sample
Model Rankings
Accuracy by Line of Business
Sample-level metrics per LOB · all table schemas
Accuracy by Document Length
Page bin = number of pages in source document
Field Category Degradation by Document Length
Category-level micro accuracy per page bin for a selected model
Model:
Accuracy by Table Schema
Each table schema evaluated independently
Show:
Field-Level Analysis
Category:
Sort:
Pricing confirmed July 2026. † Opus 4.8 cost assumed at $5/$25 MTok (Opus 5 tier).
Sonnet 5 standard rate ($3/$15); intro rate $2/$10 through Aug 31 2026.
Gemini costs include reasoning tokens from trace logs.
Pareto frontier excludes our Loss Run Model.
How to read this
- Field accuracy
- The share of fields read correctly, averaged across documents. This is the number most accuracy claims refer to. The dashboard calls it Macro Accuracy. Micro Accuracy is the same idea, weighted by document size.
- Read perfectly
- The stricter test: the share of loss runs where every field came back correct. Only these can move on without a person checking them. The dashboard calls it % Perfect.
- Cost per loss run
- What one document costs to read at each vendor's list price. Our Loss Run Model runs on our own hardware, so its cost is machine time: a single B200 GPU at $0.16633 per minute, shared across eight parallel workers.
- Time per loss run
- How long one document takes, wall clock. For an underwriter, this is turnaround on a submission.
Use the Metric switch in the sidebar to change every chart and table between the three views.
Bevaya Labs · Full writeup
The Case for Specialization
Why a specialist model beats frontier models on loss runs, how we trained it, and how each number on this page was measured.
Hunter Heidenreich and Ben Elliott · August 19, 2026 · 13 min read
Read the full writeup


