Skip to content
Bevaya Labs

Benchmarks

How we measure our models against frontier models on real insurance documents. Bevaya's models lead every frontier model tested on accuracy, at a fraction of the cost.

See the latest results
Loss run extractionhigher is better

Macro accuracy

Bevaya0%
Fable 50%
Opus 50%
Gemini 3.60%
GPT-50%
93.1%
Macro accuracy
+7.8 pts above next-best model
63×
Cost advantage
cheaper than Fable 5 at same task
2.4×
Faster
than Fable 5 · 20.9s vs 49.9s

Loss run extraction

346 loss run documents read into structured claims and policy tables. Our Loss Run Model against 12 frontier models across 18 configurations.

LatestUpdated August 2026Methodology

First published run. Frontier models shown with and without reasoning where the provider offers both. Pricing confirmed July 2026.

Loss run extraction results

Cost vs % Perfect Log scale · per sample (USD)
Latency vs % Perfect Log scale · wall-clock seconds per sample
Model Rankings
Accuracy by Line of Business Sample-level metrics per LOB · all table schemas
Accuracy by Document Length Page bin = number of pages in source document
Field Category Degradation by Document Length Category-level micro accuracy per page bin for a selected model
Model:
Accuracy by Table Schema Each table schema evaluated independently
Show:
Field-Level Analysis

How to read this

Field accuracy
The share of fields read correctly, averaged across documents. This is the number most accuracy claims refer to. The dashboard calls it Macro Accuracy. Micro Accuracy is the same idea, weighted by document size.
Read perfectly
The stricter test: the share of loss runs where every field came back correct. Only these can move on without a person checking them. The dashboard calls it % Perfect.
Cost per loss run
What one document costs to read at each vendor's list price. Our Loss Run Model runs on our own hardware, so its cost is machine time: a single B200 GPU at $0.16633 per minute, shared across eight parallel workers.
Time per loss run
How long one document takes, wall clock. For an underwriter, this is turnaround on a submission.

Use the Metric switch in the sidebar to change every chart and table between the three views.

Bevaya Labs · Full writeup
The Case for Specialization
Why a specialist model beats frontier models on loss runs, how we trained it, and how each number on this page was measured.
Hunter Heidenreich and Ben Elliott · August 19, 2026 · 13 min read
Read the full writeup