For applicants

Your insurer asked for something you don't know how to produce.

Ground truth. Error rate. A dataset of outputs against validated outcomes. These are reasonable questions and almost nobody can answer them first time. Your broker handles the insurance side — this part is model design, and it is where applications stall. Here is what the question actually means, and how to answer it.

What they are actually asking for

Two columns.

An insurer covering AI wants to know how often your model is wrong and by how much. To measure that they need two things: what your model predicted, and what the right answer turned out to be. From those two columns, the error rate is arithmetic.

That is the whole ask. Not your training data, not your source code, not your prompts. The reason applications stall is rarely that the answer is bad — it is that nobody explained the question in language a builder can act on.

In plain language

Ground truth is the correct answer, established by something other than the model.

It usually arrives one of four ways. Every one of them counts.

A person reviews and decides

Someone qualified checks a sample of the model's outputs and records the right answer. A claims handler confirms whether the flagged transaction really was fraud.

A downstream outcome tells you later

The real world files the answer after the fact. The invoice that was finally accepted shows what the freight rate actually was — you match it back to what the model said at the time.

A labeled test set

A held-back sample where the answers were established in advance, before the model saw them. If your team built one during development, that counts.

You don't have one yet

Common, and fixable. It is a starting point, not a failure — and it is far better to say so now than to have the insurer discover it at the quote. We show you the three ways operators put one in place.

A worked example

Why it is harder than it looks.

A company runs a model that scans websites for accessibility compliance. When they applied for cover, the loss they wanted insured was the demand letter from a lawyer — the moment a compliance failure turns into a legal claim.

But the letter arrives roughly once a year whether the model is right or wrong. It measures how litigious the world is, not how good the model is. It is not an AI error.

The AI error is the gap between what the model reported and what a monthly expert review of the same websites found — a miss the model made, counted against a right answer a person established. Two different things, and separating them is the whole exercise. Until you can say which losses are the model being wrong, no insurer can price the risk — and once you can, the application mostly writes itself.

What we do about it

Drafted from the application you already filed.

What your model decides, how you would know it was wrong, where the right answer comes from, what a wrong answer costs: we read those answers out of the application on file and show them back to you, each one quoting the line it came from. You correct what is wrong. You do not compose it from nothing.

At the end you have two things: a definition an underwriter can read, and the section of the insurer's form that asks about it already filled in. This is a real output for a freight operator's model:

Generated definition — freight rate quoting model
Freight Rate Quoting Engine is a prediction model used to forecast the freight rate needed to quote each shipment before it is offered to a carrier. It produces roughly 18,000 outputs a year. An output counts as an AI error when it differs from the ground truth by more than 5% of the true value. Ground truth comes from a downstream outcome observed after the fact: the final accepted carrier invoice is matched back to the quoted shipment and records the freight rate that was actually agreed and paid. Cover is sought for false negative / underprediction and false positive / overprediction. Inputs, model version and timestamp are retained, with one documented exception: customer-provided source documents cannot be retained beyond 90 days because customer contracts require deletion; the output, model version and timestamp remain available. Losses are evidenced by the customer quote, the accepted carrier invoice and the payment or settlement record, joined by shipment reference to the retained output, model version and timestamp.

Every sentence traces back to an answer the applicant gave. You can edit yours before it goes anywhere.

How the data reaches us

Three ways in. Easiest first.

  1. 1

    Connect your model and it reports on its own

    One link to whoever runs the model. They paste one thing, and that is the integration — the model reports what it predicted, and outcomes get matched as they arrive.

  2. 2

    Upload a spreadsheet you already have

    Predictions in one column, the right answer in another. If your team has been keeping any record at all, it is probably enough to start.

  3. 3

    Enter a sample by hand

    No connection, no file? Type in fifty examples. Fifty rows is enough to start — we tell you plainly which tier of evidence you have.

What you get back

Their form, their format, your standing.

Your insurer's own form, completed — their fields, their file format, filled from your evidence. And a plain statement of where you stand: what is proven, what is claimed, what is missing, and exactly what closes each gap. You see everything before anything is sent, and you decide when to send it.

What happens next

Days, not weeks.

The guided definition takes minutes. A spreadsheet upload takes as long as finding the file. A connection is one paste by whoever runs your model. Most applicants go from the insurer's question to a completed package in days, and the package stays current as new outcomes arrive — so a renewal never starts from a blank page.

Find out where you stand.

Bring the question your insurer asked. The assessment is free, the definition is yours to keep, and nothing is sent anywhere without your say-so.