Blog

Insurance-Trained vs. General-Purpose AI on Loss Runs: What Bevaya's Benchmark Found

Written by Bevaya Experts | September 23, 2026

Newer general-purpose AI models are not closing the accuracy gap on insurance documents. In Bevaya's latest benchmark of 346 real loss runs, every general-purpose model tested scored between 78% and 85% field accuracy, regardless of how new or expensive it was. Bevaya's insurance-trained InsurGPTâ„¢ loss run model scored 93.1%.

With the attention general-purpose AI has drawn, a fair question follows: is it good enough for insurance work on its own? The benchmark measured that on the industry's hardest document and published the result.

The test compared Bevaya's model with 12 leading general-purpose AI models on the same loss runs, under the same scoring, with the general-purpose models used the way most organizations would use them: through their APIs, with an optimized prompt and no insurance-specific training. The full method and every result are on the benchmarks page. This article covers what the test found and what it means for an insurer bringing AI into underwriting and claims. 

Key Findings 

  • Bevaya's InsurGPT loss run model led every accuracy measure, scoring 93.1% field accuracy against 85.3% for the strongest general-purpose model.

  • Every general-purpose model landed between 78% and 85%, across four model families and three model generations, from the most expensive to the least.

  • Bevaya's model returned a loss run needing no correction nearly twice as often as the strongest general-purpose model, and read each document in less than half the time.

  • Newer general-purpose models scored within a point or two of the ones they replaced. The difference came from training on insurance documents, not from a larger or newer model. 

 

Why the Gap Is Bigger Than It Looks  

Eight points of field accuracy may sound like a small difference. On a loss run carrying 100 values, 93.1% leaves about seven incorrect fields and 85.3% leaves about fifteen, more than twice as many. And a loss run is usable only when every value on it is correct. One wrong reserve figure or claim number means the whole file needs a person, which is why the general-purpose models returned a document needing no correction roughly half as often as Bevaya's model did.

The gap widened on loss runs running past ten pages. More pages means more fields, and each field is another chance for one to be wrong. For an underwriting or claims team, the number that matters is not how many fields came back right but how many files still have to be opened.

 

Why Insurance Documents Are Different

Loss runs were chosen for the benchmark because they are among the hardest documents in insurance to read. Every carrier formats its own differently. Total incurred, the figure underwriters rely on most, is often not printed and has to be calculated from paid loss, paid expense and reserves in other columns.

Claim numbers are arbitrary strings that sit in a different place on every form. None of that can be worked out from the page, and because loss runs are private, there is no public source a general-purpose model could have learned it from. That is why a newer or larger general model does not close the gap on its own.

InsurGPT, Bevaya's ensemble of specialized AI models, is trained on more than 300 million non-public insurance documents labeled by what Bevaya calls model tutors: insurance practitioners on staff whose sole job is creating training data. The loss run model alone required seven months of their annotation.

 

What This Means When Evaluating AI

Whether an insurer is weighing a vendor or considering building on a general-purpose model directly, the question is the same: what does that model do on a real insurance document, and what changes when it has been trained on insurance documents.

Most AI products sold into insurance are built on general-purpose models accessed through an API, and they inherit whatever the underlying model can do. Workflow and integrations add value on top, but they do not move the model out of the range the test found.

Two questions serve any evaluation. Which model reads the document, and was it trained on insurance documents or accessed through a general-purpose API. And what is the published evidence for its accuracy, on how many documents, scored by what method. Prompting alone rarely gets an AI agent to production, which is a big part of why the build versus buy decision matters.

 

How the Benchmark Relates to Production Accuracy

The benchmark measures the model on its own, with no verification, no confidence thresholds, and no human review, so that the comparison isolates the model.

Production works differently. Every Bevaya prediction passes through the Verifier model, which checks it against the source document, confidence is scored field by field, and anything uncertain is routed to the insurer's own staff through Bevaya's patented human-in-the-loop technology, with every action recorded to a full audit trail.

That is what carries accuracy from 93.1% in the benchmark to the 98%+ Bevaya delivers in production, where more than 70% of documents are completed with no human touch. The two figures measure different things. The benchmark is the harder question underneath the production number, and a model that starts at 93.1% gives the system around it less to catch than one that starts at 85%.

 

Insurance Is Different, and the Model Has to Be Too

General-purpose AI is a remarkable technology, and on general work it keeps getting better. Insurance work is not general. A loss run carries conventions that belong to one carrier, financial columns that mean something only to someone who knows how reserves and paid losses relate, and identifiers that follow no pattern at all.

That knowledge lives in private documents and in the people who read them, and it is not in any public dataset a general model can learn from. No amount of scale reaches it, and the benchmark shows what that looks like in practice: every general-purpose model in the same band, whatever its generation or price, and the insurance-trained model ahead on every measure.

 

 

That is the conclusion an insurer should carry into any AI evaluation. A model built for everything was not built for this, and a newer release does not change what the model was trained on. Accuracy on insurance documents comes from training on insurance documents, by people who understand them, with the checks and human review around the AI model that let the work go through at 98%+.