Blog

AI for Loss Runs: What Bevaya's Benchmark Means for Underwriting Teams

Written by Bevaya Experts | September 23, 2026

A commercial submission rarely arrives with one loss run. It arrives with three or four, from different carriers, in different formats, and someone has to turn them into a single picture of prior loss history before an underwriter can look at the risk. That reconciliation is where submission turnaround goes.

Bevaya builds AI Agents that do that reading. A loss run arrives, the AI Agent pulls every value off it into the structured fields your systems expect, checks its own work, and puts anything it is unsure of in front of a person before it goes anywhere.

Underneath each AI Agent is InsurGPTâ„¢, Bevaya's set of AI models trained on insurance documents, including one built specifically for loss runs. The question underwriting leaders reasonably ask is whether that AI is accurate enough to be trusted with the work, and whether the general-purpose AI tools in the news could do it just as well.

Bevaya published a benchmark to answer both. It compared the InsurGPT loss run model with 12 of the leading general-purpose models on 346 real loss runs, with every model reading the same documents and scored the same way. 

 

How Many Loss Runs Actually Come Back Ready to Use 

Most AI accuracy claims count individual values. A loss run holds a hundred or more of them, and an underwriting desk does not work in values. One wrong total incurred or one misplaced policy number sends the whole document back to a person, and the time saved on everything else is gone. That is what the benchmark measures.

Bevaya's InsurGPT loss run model made less than half as many errors as the strongest general-purpose model in the test, which meant nearly twice as many loss runs came back needing no correction at all. The gap widened on loss runs running past ten pages. 

On an underwriting desk that is the difference between a loss history that is ready when the underwriter opens the submission and one that waits for someone to reconcile it. Underwriting judgment goes to the risk instead of to four carriers' formats, and quotes go back faster. 

 

The Two Values on a Loss Run Underwriters Cannot Afford to Get Wrong  

A loss run can hold dozens of fields, but a mistake in either of these two changes the underwriting decision itself.

  • The policy number attaches a loss to the right risk. If it is wrong, the loss history an underwriter is pricing from belongs to another account, and nothing downstream catches it because every other value looks fine.

  • Total incurred is the figure underwriters rely on most, and on many loss runs it is not printed anywhere. It has to be worked out from paid loss, paid expense, and outstanding reserves sitting in separate columns. A general-purpose model has no way to know that, so it finds a number that looks close or leaves the field empty. Bevaya's model calculates it, because it was trained on the documents where that convention lives.

That is the difference the benchmark measured. Every carrier formats loss runs its own way, and those conventions exist only inside private insurance documents. Bevaya's AI models are trained on hundreds of millions of non-public insurance documents, annotated by insurance practitioners. There is nothing public for a general-purpose model to have learned this from, which is why the newest general models only scored within a point or two of the ones they replaced.

 

What Happens When the AI Is Not Sure

Bevaya's benchmark measured the InsurGPT loss run model on its own. In a live deployment, it does not work alone. Every value it reads carries a confidence score, and anything below the threshold your team sets goes to your own reviewer before it reaches a submission or policy system. A second AI model, the Verifier, checks each value against the source document, and every action is recorded, so an auditor can see where a value came from and who approved it.

That combination is what takes accuracy to 98%+ in production, with more than 90% of loss runs completed without anyone touching them. A reviewer sees only the handful of values the system flagged, not the whole document again. 

 

Two Questions Worth Asking Any AI Vendor

Every vendor will tell you their model performs well, and these two questions show you what that claim is built on.

  • How often does the whole loss run come back with nothing to fix? Most vendors answer with an accuracy percentage, which counts individual values. That number can look excellent while your team still opens every file, because one wrong value is enough to send a document back. The share of documents that need no correction is the only figure that tells you how much work actually lands on a person.

  • How does that hold up on a 20-page loss run? Accuracy on a two-page loss run tells you little about the submissions that matter. Longer documents hold more values, so the chance that all of them are right falls as the document grows. Ask for the result on the lengths your team sees.

A vendor who can answer both has measured the thing that predicts your throughput. One who can only quote a single accuracy figure has measured something else.

 

Bevaya's answers are published on the benchmarks page, model by model and length by length.