How the lab was built, measured and tested

How the numbers were made, and how far to trust each one.

Figures last updated , from the published evidence. What changed

From synthetic claims to a published result, in eleven steps

The decision model is asked once and its answers are stored, so every workflow can be tried again for free. Only the final test asks it afresh, once, on claims no one has seen. Swapping either model is one setting, from a reviewed list, and every model goes through the same steps.

Once, paid

  1. Synthetic claimsA seeded generator: 1,000 claims with known right answers; frauds, manipulation and missing papers planted
  2. The decision model answersJev, through OpenRouter's decisions endpoint, in TypeSafe's System One format. It could as well be TypeSafe direct, a free practice model, or a language model wrapped to answer in the same format
  3. Every paid call approvedA ledger approves each paid run in advance, within a fixed budget for the whole lab
  4. Answers storedEvery answer and probability kept, with the model and question versions

Free, as often as needed

  1. Replay every workflowEvery workflow reads the same stored answers: only the decision changes
  2. Cross-validationEach claim is scored by settings tuned on the other claims, never on itself
  3. AdmissionA design earns a real test only if it passes on at least nine in ten of a hundred different splits of the claims
  4. FreezeThe weights and the cut-off are written down and never changed again

One shot, paid

  1. Test on fresh claims1,000 new claims from a seed drawn when the run starts; the decision model asked afresh; the result final
  2. ConfirmationFor a version that passed: 2,000 more, judged on the bottom of the accuracy range
  3. Published evidenceExported byte for byte with its hashes; every figure on this site traces to it

The lever. Accuracy and speed move with the model and with the process around it. A new model goes through all eleven steps. A new question to the decision model, or a new step such as asking the customer, goes through the same tests as every workflow before it: nothing reaches this site until it has faced fresh claims.

The decision model, measured on this run

Every call the run made is recorded, with its time and its cost. Its yes-or-no answers were checked against the known right answers: the chart puts the probability it gave beside how often the answer was yes.

Calls
13,278 on 1,100 claims, every test on this run
Time a call
276 ms median; one in twenty over 515 ms
Cost
$0.40 in all: $0.03 per thousand calls
Calibration
9.6 points off on average, over 7,000 yes-or-no answers
0%0% 50%50% 100%100% what it said how often yes
Each dot is a tenth of its yes-or-no answers, grouped by the probability it gave: across, the average probability; up, how often the answer was yes. On the diagonal the two agree. At the top the answer was yes more often than it said, and at the bottom less often: mostly more cautious than it needed to be.
The ten groups
AnswersWhat it said, on averageHow often yes
7003.1%0.0%
7005.2%0.6%
7007.5%0.6%
70016.9%4.1%
70035.7%30.1%
70057.9%27.0%
70077.5%84.3%
70083.3%93.6%
70087.6%97.6%
70095.3%100.0%

What the synthetic claims contain, and how a decision is judged right

Every figure on this site comes from synthetic claims the lab generated. Every set of a thousand is built to the same quotas, so the fresh sets used in the tests can be set beside the claims the designs were developed on. Categories are counted here; how a fraud or a manipulation attempt is built is not described.

What each set of claims is made of, as shares of its claims
Share of claimsDevelopment claims (1,000)Fresh claims, pooled (10,000)
By line of insurance
Auto40.0%40.0%
Home30.0%30.0%
Commercial20.0%20.0%
Life10.0%10.0%
By amount claimed
Under $20,00075.5%75.4%
$20,000 to $50,00016.5%16.6%
Over $50,0008.0%8.0%
Planted by the generator
Fraud to investigate10.0%10.0%
Manipulation attempts5.0%5.0%
Ambiguous15.0%15.0%
Missing documents5.1%4.5%
Who must see it
Must reach a person, by rule21.6%21.2%
The most any workflow could settle with no person78.4%78.8%

The planted shares are the lab's own quotas, not any insurer's mix of claims. Accuracy on the claims settled with no person depends on that mix, and would differ with another.

How a decision is judged right

Every synthetic claim carries an answer key written by the generator from the facts it planted: the dispositions it accepts. A decision is judged against that key, never against a person's or a model's opinion.

Paid with no person
Paying without a person is right when the claim's answer key accepts payment.
Asked for more, with no person
Asking the customer for more is right when the answer key accepts a request for more information, as it does for a claim missing documents.
Sent to a person
Sending the claim to a person is right when the answer key calls for review, investigation or quarantine. An ambiguous claim's key always accepts a person.
Wrong
A settled claim is wrong when its disposition is not one its answer key accepts. Claims sent to a person are not settled and are not counted as wrong.

The routing rule. A claim over $50,000, a claim the answer key marks as fraud to investigate, and a manipulation attempt to quarantine must always reach a person. Their share caps what any workflow can settle with no person.

How the claims are generated

Quotas are exact per 1,000 claims: by line, by category (clean, ambiguous, fraud, manipulation) and the share over $50,000. The seed changes which claims are generated, their amounts, documents and wording, never the quotas. Generator version v1.

  • The evaluation corpus is generated from one fixed, public seed; every design is developed and cross-validated on it, so its figures are development evidence.
  • A held-out test's seed does not exist until its run starts, from a slot declared in advance for one frozen version. The tests run before 28 September drew it from the operating system's random source at that moment; from the declaration log v2 on, it is computed from a public drand round the declaration named a day ahead.
  • A confirmation adds a second 1,000 claims from a seed derived from the first. Held-out corpora carry no hand-written fixtures and no rewordings, and no held-out seed is reused in any pool.
  • A test declared in advance for the version as it runs now, payment limit included, runs once on 3,000 claims: from its slot's seed and two seeds derived from it, each checked against every earlier seed and registered to that slot.

The process, today and new

An adjuster reads every claim, checks the policy by hand and decides. It is careful and slow: most of the waiting is a claim sitting in a queue.

Of 6 steps, people do 5, code 1, the decision model none.
  • Code
  • Decision model
  • People
  1. Code: Claim arrives
  2. People: Waits in a queue
  3. People: Adjuster reads everything
  4. People: Adjuster checks the policy by hand
  5. People: Adjuster decides
  6. People: Payment is set up

Where the 1,000 claims end up:

All 1,000 decided by a person, at step 6
What each step does
  1. Claim arrives
  2. Waits in a queue Most of the 5 days before a decision (an assumed wait) is spent here.
  3. Adjuster reads everything
  4. Adjuster checks the policy by hand
  5. Adjuster decides
  6. Payment is set up

By hand, a claim takes about 122.0 min of staff time and waits about 5.0 days for a decision, at benchmark staff times and the lab's assumed wait.

  1. Code: Claim arrives
  2. Decision model: Jev answers narrow questions
  3. Code, a fixed rule: Over $50,000?
  4. Code, a fixed rule: Manipulation attempt?
  5. Code, a fixed rule: Fraud signs: how strong?
  6. Code, a fixed rule: Policy in force?
  7. Code: Certain enough for the amount?
  8. Code: Over $20,000?
  9. People: Quick confirm, requests only
  10. Code: Pay, or ask for what is missing
  11. People: Audit, afterwards

Where claims go:

a person, with the analysis attached, leaving at step 3
quarantine, leaving at step 4
a light check or an investigation, leaving at step 5
a person, who confirms any denial, leaving at step 6
a standard review by a person, leaving at step 7
a person approves the payment, leaving at step 8
paid, or asked for documents, with no person, at step 10
What each step does
  1. Claim arrives Code notes the amount and looks up the policy once. No model has been asked anything yet.
  2. Jev answers narrow questions Fourteen typed questions about completeness, evidence and fraud, each answered with a probability. Jev answers; it never decides.
  3. Over $50,000? Fixed in code. No setting and no answer from the decision model can change it.
  4. Manipulation attempt? Code acts on Jev's answer about hidden instructions in the text.
  5. Fraud signs: how strong? Weak signs get a light check by a person first, and a claim it clears gets a standard review. Strong signs, or a failed check, go to a full investigation.
  6. Policy in force? One lookup in the policy system. A person confirms every denial; none is automatic.
  7. Certain enough for the amount? W3 weighs Jev's answers. Above $5,000 the doubt allowed shrinks as the amount grows, and it settles a claim alone only if one minus the doubt clears 0.813.
  8. Over $20,000? Nothing over the limit is paid automatically, whichever step proposes it: a person approves it.
  9. Quick confirm, requests only A request for more information that code is nearly sure of gets one quick look by a person. Every approval a person sees gets a full review.
  10. Pay, or ask for what is missing In about a second, with every step recorded.
  11. Audit, afterwards A random sample of the claims settled with no person is checked by a person afterwards.
The fixed rules, as the engine records them
  • Claims over $50,000 require mandatory human review
  • Suspected manipulation of automated processing quarantines the claim
  • Claims with fraud indicators require investigation
  • Claims on a policy that is not in force go to a human (denials are never automatic)

Every version on one table

Development figures on the 1,000 claims the lab was built with. Rows marked + keep W3's decision and change what happens around it. Settled alone counts paid and asked for documents together: the evidence does not yet show them apart for development. Each version's detail follows. Every row here uses the first way of choosing the cut-off, which failed its test on fresh claims, with no payment limit; the recommended design below uses the corrected way of choosing the cut-off, as in the version that passed, with the payment limit.

VersionSettled aloneRightStaff time per claimNet per 1,000Test on fresh claims
W0 · Trust the one big answer3.4%70.59%207.8 min−$100,030not tested
W1 · Combine the small answers36.6%100.00%187.9 min−$76,790not tested
W2 · Learn how much to trust each answer57.8%99.83%113.8 min−$39,441not tested
W3 · Be stricter as amounts rise62.6%99.68%110.9 min+$8,809Passed corrected, twice
+ A quick confirmation for near-certain claims (now for requests only)62.6%99.68%102.2 min+$18,931not tested
+ Grade the fraud response62.6%99.68%74.1 min+$51,761not tested
+ A second vote on near-misses (a negative result)63.0%99.52%74.0 min+$51,845not tested
+ Ask the customer one question66.3%99.40%72.9 min+$3,649as W4: Failed
+ Watch after the decision62.6%99.68%74.8 min+$50,909not tested
Claims settled with no person, paid and asked together share of all claims
0%20%40%60%80%HandHand Today, by hand: 0.0%W0W0 Trust the one big answer: 3.4%W1W1 Combine the small answers: 36.6%W2W2 Learn how much to trust each answer: 57.8%W3W3 Be stricter as amounts rise: 62.6%+QC+QC A quick confirmation for near-certain claims (now for requests only): 62.6%+Fraud+Fraud Grade the fraud response: 62.6%+2nd+2nd A second vote on near-misses (a negative result): 63.0%+Ask+Ask Ask the customer one question: 66.3%+Audit+Audit Watch after the decision: 62.6%most possible 78.4%
Staff minutes per claim lower is better
0100200300HandHand Today, by hand: 122.0 minW0W0 Trust the one big answer: 207.8 minW1W1 Combine the small answers: 187.9 minW2W2 Learn how much to trust each answer: 113.8 minW3W3 Be stricter as amounts rise: 110.9 min+QC+QC A quick confirmation for near-certain claims (now for requests only): 102.2 min+Fraud+Fraud Grade the fraud response: 74.1 min+2nd+2nd A second vote on near-misses (a negative result): 74.0 min+Ask+Ask Ask the customer one question: 72.9 min+Audit+Audit Watch after the decision: 74.8 minby hand 122.0 min
Days until the claimant has a decision lower is better
0246HandHand Today, by hand: 5.0 daysW0W0 Trust the one big answer: 4.8 daysW1W1 Combine the small answers: 3.2 daysW2W2 Learn how much to trust each answer: 2.1 daysW3W3 Be stricter as amounts rise: 1.9 days+QC+QC A quick confirmation for near-certain claims (now for requests only): 1.4 days+Fraud+Fraud Grade the fraud response: 1.4 days+2nd+2nd A second vote on near-misses (a negative result): 1.4 days+Ask+Ask Ask the customer one question: 1.6 days+Audit+Audit Watch after the decision: 1.4 days
Net per 1,000 claims against doing it all by hand at benchmark staff costs
−$150k−$100k−$50k$0+$50k+$100kHandHand Today, by hand: $0W0W0 Trust the one big answer: −$100,030W1W1 Combine the small answers: −$76,790W2W2 Learn how much to trust each answer: −$39,441W3W3 Be stricter as amounts rise: +$8,809+QC+QC A quick confirmation for near-certain claims (now for requests only): +$18,931+Fraud+Fraud Grade the fraud response: +$51,761+2nd+2nd A second vote on near-misses (a negative result): +$51,845+Ask+Ask Ask the customer one question: +$3,649+Audit+Audit Watch after the decision: +$50,909
The figures behind the charts
VersionClaims settled with no person, paid and asked togetherStaff minutes per claimDays until the claimant has a decisionNet per 1,000 claims against doing it all by hand
Hand Today, by hand0.0%122.0 min5.0 days$0
W0 Trust the one big answer3.4%207.8 min4.8 days−$100,030
W1 Combine the small answers36.6%187.9 min3.2 days−$76,790
W2 Learn how much to trust each answer57.8%113.8 min2.1 days−$39,441
W3 Be stricter as amounts rise62.6%110.9 min1.9 days+$8,809
+QC A quick confirmation for near-certain claims (now for requests only)62.6%102.2 min1.4 days+$18,931
+Fraud Grade the fraud response62.6%74.1 min1.4 days+$51,761
+2nd A second vote on near-misses (a negative result)63.0%74.0 min1.4 days+$51,845
+Ask Ask the customer one question66.3%72.9 min1.6 days+$3,649
+Audit Watch after the decision62.6%74.8 min1.4 days+$50,909

W0 to W3: turning the answers into a decision

All four change one box in the process: is the evidence strong enough to act alone? The same answers from the same model are read four different ways.

W0 Trust the one big answer Measured in the engine Lab verdict: Abandoned: not accurate enough and loses money

The engine acts on its own only when the model confidently answers one broad question, "what do you recommend?", and every hand-set guard agrees.

What changed: The engine as shipped, hand-set 0.8 confidence threshold, every claim, no cross-validation. The model's answers now reach the engine. The engine, its rules and its hand-set thresholds are exactly as first specified.

Settled with no person (paid and asked together)
3.4%
Wrong decisions
10 of 34 settled alone
Staff time per claim
207.8 min
Days to a decision
4.8 days
Wrongly paid
$0
Net per 1,000 claims
−$100,030 against doing it all by hand
  • Settled with no person 34
  • Full review 545
  • Specialist 244
  • Fraud investigation 145
  • Quarantine 32

What improved

  • Handled with no person: 0.0% to 3.4%.
  • Average wait for a decision: 5.0 to 4.8 days.

Risks and open issues

  • Staff time per claim: 122.0 minutes to 207.8 minutes.
  • All 10 errors are valid claims wrongly asked for more information. Accuracy on handled claims is 70.6% (24 of 34), below the 99.5% bar, with $0 wrongly paid. That is why the first design was abandoned.
  • Almost every claim still reaches a person, so the model adds cost without taking work away.
  • net per 1,000 claims is $-100,030, not above $0
One claim: Auto claim for $1,850 C-0001
Before
By hand: waits about 5 days
In this step
Sent to a human adjuster, because the model was not confident enough in its one broad answer

Whatever happens to the simplest claim shows what this version can and cannot do.

W1 Combine the small answers Cross-validated Lab verdict: Improved, does not pass: accurate but loses money

Ignore the broad recommendation. Code combines the model's narrow answers (complete? photos match? fraud risk?) and acts on the weakest one.

What changed: How answers are combined. No new questions, no new model calls.

Settled with no person (paid and asked together)
36.6%
Wrong decisions
0 of 366 settled alone
Staff time per claim
187.9 min
Days to a decision
3.2 days
Wrongly paid
$0
Net per 1,000 claims
−$76,790 against doing it all by hand
  • Settled with no person 366
  • Full review 213
  • Specialist 244
  • Fraud investigation 145
  • Quarantine 32

What improved

  • Handled with no person: 3.4% to 36.6%. The two were measured differently: chapter 1 is the engine as shipped, run once over every claim; this chapter is the lab's cross-validated score over the same answers, so part of the jump is the method, not only the design.
  • Staff time per claim: 207.8 minutes to 187.9 minutes.
  • Average wait for a decision: 4.8 to 3.2 days.

Risks and open issues

  • Accuracy measured at 100.00%, but on 366 handled claims the honest range is 99.0% to 100.0%, so 99.5% is met but not yet proven.
  • One under-confident answer blocks a good claim; one over-confident answer could pass a bad one.
  • net per 1,000 claims is $-76,790, not above $0
One claim: Commercial claim for $18,605 C-0010
Before
Customer asked for the missing information at once, with no person (this was the wrong call)
In this step
Sent to a human adjuster, because it was not certain enough for the amount at stake

The previous version got this claim wrong. This one does not.

W2 Learn how much to trust each answer Cross-validated Lab verdict: Improved, does not pass: accurate but loses money

The model's probabilities do not match how often it is right (mostly too cautious), and some answers matter more than others. A small statistical model, fitted on past claims, learns how to weigh them.

What changed: Replaces hand-set combination rules with 14 learned weights. Still code: no free text, and the hard rules run first.

Settled with no person (paid and asked together)
57.8%
Wrong decisions
1 of 578 settled alone
Staff time per claim
113.8 min
Days to a decision
2.1 days
Wrongly paid
$49,101
Net per 1,000 claims
−$39,441 against doing it all by hand
  • Settled with no person 578
  • Full review 245
  • Fraud investigation 145
  • Quarantine 32

What improved

  • Handled with no person: 36.6% to 57.8%.
  • Staff time per claim: 187.9 minutes to 113.8 minutes.
  • Average wait for a decision: 3.2 to 2.1 days.

Risks and open issues

  • Wrongly paid out: $0 to $49,101 (1 wrong decision in 578 handled).
  • Accuracy measured at 99.83%, but on 578 handled claims the honest range is 99.0% to 100.0%, so 99.5% is met but not yet proven.
  • Weights are learned from synthetic claims and would need re-fitting on real ones.
  • net per 1,000 claims is $-39,441, not above $0
One claim: Home claim for $18,400 C-0002
Before
Sent to a specialist, because it was flagged as needing a specialist
In this step
Paid in about a second, with no person

The previous version sent this claim to a person. This one handles it correctly.

W3 Be stricter as amounts rise Cross-validated Lab verdict: Candidate

Same learned weights as W2, but a large claim needs much more certainty than a small one before it is paid without a person.

What changed: The confidence needed now rises with the claim amount. No new model calls.

Settled with no person (paid and asked together)
62.6%
Wrong decisions
2 of 626 settled alone
Staff time per claim
110.9 min
Days to a decision
1.9 days
Wrongly paid
$4,211
Net per 1,000 claims
+$8,809 against doing it all by hand
  • Settled with no person 626
  • Full review 197
  • Fraud investigation 145
  • Quarantine 32

What improved

  • Handled with no person: 57.8% to 62.6%.
  • Staff time per claim: 113.8 minutes to 110.9 minutes.
  • Average wait for a decision: 2.1 to 1.9 days.
  • Wrongly paid out: $49,101 to $4,211 (2 wrong decisions in 626 handled).

Risks and open issues

  • Accuracy measured at 99.68%, but on 626 handled claims the honest range is 98.8% to 99.9%, so 99.5% is met but not yet proven.
  • Large honest claims wait for a person more often. That is the intended trade.

This step was frozen and tested on fresh claims twice. The first version failed the accuracy bar; a second, with a stricter cut-off, passed. It then passed a larger, stricter confirmation. The figures here use the first version's cut-off. The held-out tests.

One claim: Life claim for $49,101 C-0476
Before
Paid in about a second, with no person (this was the wrong call)
In this step
Sent to a human adjuster, because it was not certain enough for the amount at stake

The previous version got this claim wrong. This one does not.

The lesson. Accuracy is the wrong single measure. A version can be right almost every time and still lose money, because one large wrong payment outweighs many small savings. What matters is how likely a mistake is, times how much it costs.

Around W3: the work people still do, and what else was tried

These keep W3's decision and change what happens around it: a short confirmation instead of a full review, a light check before a fraud investigation, a language model's second opinion, a question to the customer, and an audit afterwards.

+ A quick confirmation for near-certain claims (now for requests only) Replayed, assumed minutes Lab verdict: Kept

A claim the decision model has fully prepared, but is not quite sure enough to pay, does not need a full review. It needs one look at the one doubtful item.

What changed: Adds a lane between "no person" and "full review". It does not raise the no-person rate; it cuts the minutes spent on claims that still reach someone.

Settled with no person (paid and asked together)
62.6%
Wrong decisions
2 of 626 settled alone
Staff time per claim
102.2 min
Days to a decision
1.4 days
Wrongly paid
$4,211
Net per 1,000 claims
+$18,931 against doing it all by hand
  • Settled with no person 626
  • Quick confirmation 168
  • Full review 29
  • Fraud investigation 145
  • Quarantine 32

What improved

  • Staff time per claim: 110.9 minutes to 102.2 minutes.
  • Average wait for a decision: 1.9 to 1.4 days.
  • The no-person rate, accuracy and wrongly-paid figures are unchanged by design: this chapter only changes what happens to claims that still reach someone.

Risks and open issues

  • 60 of 168 quick confirms were overridden because the prepared suggestion was not right for the claim. Each one pays for the quick look and the full work.
  • A quick confirm is assumed to take 3 minutes and wait 1 day. Both are placeholders, not measurements.
  • The replay assumes the person always catches a wrong suggestion. People may approve by habit when the file looks finished.

Research on people working with automated decisions predicts approval by habit: on decision tasks, a person and an AI together did worse on average than the better of the two alone, and people miss a wrong answer most when it is rare. The replay cannot show whether real reviewers would catch a wrong prepared file. Testing reviewers with planted wrong files is a requirement of this lane, not an option.

In a simulated reviewer, with assumed costs on the development claims, a quick confirmation beats a full review only if the reviewer catches at least 99.6% of wrong files: 51 of the 110 prepared files are wrong, 38 of them approvals worth $977,936. People reviewing automated work miss far more than that. For requests for more information, where a missed wrong file costs only a delay, it pays at any catch rate: 28 requests, 13 of them wrong. So the recommended design keeps the quick confirmation for requests only.

One claim: Auto claim for $844 C-0626
Before
Sent to a human adjuster, because it was not certain enough for the amount at stake
In this step
A person confirms the prepared file in about 3 minutes (assumed)

This claim could have been handled automatically and missed the bar narrowly.

+ Grade the fraud response Replayed, assumed minutes Lab verdict: Kept

Today one yes-or-no test sends a claim to a full investigation. Weak signals deserve a light check first, with the evidence file assembled by the decision model either way.

What changed: Replaces the single fraud exit with two levels, and hands investigators a prepared file.

Settled with no person (paid and asked together)
62.6%
Wrong decisions
2 of 626 settled alone
Staff time per claim
74.1 min
Days to a decision
1.4 days
Wrongly paid
$4,211
Net per 1,000 claims
+$51,761 against doing it all by hand
  • Settled with no person 626
  • Quick confirmation 168
  • Light fraud check 78
  • Full review 29
  • Fraud investigation 67
  • Quarantine 32

What improved

  • Staff time per claim: 102.2 minutes to 74.1 minutes.
  • The no-person rate, accuracy and wrongly-paid figures are unchanged by design: this chapter only changes what happens to claims that still reach someone.

Risks and open issues

  • 30 claims where an investigation was warranted met the light check first. The replay assumes that check catches them; if it does not, this is where fraud gets through.
  • A light check is assumed to take 20 minutes. It is a placeholder, not a measurement.
  • No saving is claimed for handing investigators a prepared file; that would need real timings.

This step models the light check as perfect: every claim that warrants an investigation goes on to one, and every needless referral is cleared by the check alone. In a simulation with assumed rates on the development claims, giving every cleared claim its standard review, the graded check would be worth +$28,980 per thousand claims over no graded check if it missed nothing. It pays only if it catches at least 72.7% of the 23 claims among its 78 that warrant an investigation, or if the standard review after it catches what it misses, in which case it pays at any catch rate. Catching 50% with no backstop, it comes to −$24,079 per thousand claims. The recommended design gives every cleared claim its standard review, never an automatic decision, and its figures count it.

One claim: Home claim for $7,200 C-0005
Before
Sent to fraud investigation, because fraud indicators were above the limit
In this step
A light verification clears it and it rejoins the normal flow

An honest claim that tripped a hand-set fraud limit.

+ A second vote on near-misses (a negative result) Tested with real calls Lab verdict: Kept

Jev is fast and cheap but cautious. For the few claims just under the bar, a language model reads the claim and seven of Jev's findings; agreement, above a certainty floor fitted on other claims, lifts the claim over.

What changed: Adds one model call, for near-misses only, so the cost stays small.

Settled with no person (paid and asked together)
63.0%
Wrong decisions
3 of 630 settled alone
Staff time per claim
74.0 min
Days to a decision
1.4 days
Wrongly paid
$4,211
Net per 1,000 claims
+$51,845 against doing it all by hand
  • Settled with no person 630
  • Quick confirmation 164
  • Light fraud check 78
  • Full review 29
  • Fraud investigation 67
  • Quarantine 32

What improved

  • Handled with no person: 62.6% to 63.0%.
  • Staff time per claim: 74.1 minutes to 74.0 minutes.
  • Tested with real calls to a second model (anthropic/claude-haiku-4.5) on 110 claims, for $0.09 in total.

Risks and open issues

  • Accuracy measured at 99.52%, but on 630 handled claims the honest range is 98.6% to 99.8%, so 99.5% is met but not yet proven.
  • Agreement alone is unsafe. With no certainty floor, the second model agreed on 13 claims and 6 were wrong, paying out $148,084. The floor, fitted on other claims, is what keeps those with a person.
  • 1 agreed claim was wrong, paying out $0: two models can be wrong together. Agreement is evidence, not proof.
  • Only 4 claims were added, so this chapter's own error rate is known very loosely.
  • Adds a second supplier, and sends claim text to it, for every claim it touches.

The lab's rule kept this step because it cleared the accuracy bar and did not lose money. The rule weighs accuracy and money only. It does not count a second supplier, or claim text leaving the company, and on those this step does not earn its place.

One claim: Home claim for $8,725 C-0018
Before
Sent to a human adjuster, because it was not certain enough for the amount at stake
In this step
A second model reads the claim and Jev's findings; if both agree and the claim clears the fitted floor, it is handled

Safe to handle, but Jev alone was not sure enough.

+ Ask the customer one question Simulated replies Lab verdict: Rejected: not accurate enough

Missing information is the largest fixable reason a claim cannot be decided. One specific question, answered by the customer, can turn an unclear claim into a clear one.

What changed: The only proposal that can raise the ceiling on the no-person rate, because it changes what is known about the claim.

Settled with no person (paid and asked together)
66.3%
Wrong decisions
4 of 663 settled alone
Staff time per claim
72.9 min
Days to a decision
1.6 days
Wrongly paid
$53,674
Net per 1,000 claims
+$3,649 against doing it all by hand
  • Settled with no person 663
  • Quick confirmation 102
  • Light fraud check 90
  • Full review 29
  • Fraud investigation 67
  • Quarantine 49

What improved

  • Handled with no person: 63.0% to 66.3%.
  • Staff time per claim: 74.0 minutes to 72.9 minutes.
  • All 17 manipulation attempts hidden in replies were caught and quarantined.
  • Customer replies were synthesised by a language model (anthropic/claude-haiku-4.5), by design 66 helpful, 26 unhelpful and 17 attempts to manipulate the system. Jev then re-analysed each enlarged claim with real calls. $0.21 in total.

Risks and open issues

  • Average wait for a decision: 1.4 to 1.6 days.
  • Wrongly paid out: $4,211 to $53,674 (4 wrong decisions in 663 handled).
  • Accuracy on handled claims is 99.40%, below the 99.5% bar.
  • 1 planted fraud with a reassuring reply was paid: $49,463. With the helpful-sounding reply added, the claim cleared the certainty bar. Nothing in this flow checks that what a reply describes is real.
  • 1 claim was handled wrongly after the reply, paying out $49,463. A reply is new untrusted text; planted frauds and manipulation attempts were left in the test on purpose.
  • Code allows one question only: a claim that would need a second question goes to a person.
  • The replies are invented. This shows what the flow can do with a given reply, not how real customers answer, or whether they answer at all.
  • A reply is assumed to take 2 days. Claims that still reach a person wait for the reply and then for the queue.

Every manipulation attempt was quarantined, in the claims and in the replies, but all were built from five fixed test phrasings. A defence measured on fixed attacks can look stronger than it is; subtler attacks are not yet tested.

One claim: Life claim for $49,463 C-0277
Before
Sent to a human adjuster, because it was not certain enough for the amount at stake
In this step
Paid $49,463 incorrectly after a reassuring synthesised reply.

More convincing text is not verified evidence. This planted fraud crossed the payment threshold after a reassuring synthesised reply; the flow did not verify that the evidence described in the reply was real.

+ Watch after the decision Replayed, assumed minutes Lab verdict: Kept

Random human audits can gather evidence about automated decisions; they do not automatically prove production accuracy. This demonstration examines sampling, fairness and drift within a synthetic run.

What changed: Adds a monitoring design to evaluate further. Production validation needs representative data, reliable human checks and evidence over time; complaint tracking is not built. It watches chapter 6, the recommended version: the second opinion (chapter 7) is left out by design, and chapter 8 is not built on.

Settled with no person (paid and asked together)
62.6%
Wrong decisions
2 of 626 settled alone
Staff time per claim
74.8 min
Days to a decision
1.4 days
Wrongly paid
$4,211
Net per 1,000 claims
+$50,909 against doing it all by hand
  • Settled with no person 626
  • Quick confirmation 168
  • Light fraud check 78
  • Full review 29
  • Fraud investigation 67
  • Quarantine 32

What improved

  • A 10% random audit checks 73 no-person decisions out of 626. Once 765 audited decisions in a row are clean, "at least 99.5% accurate" is supported by evidence: about 7,650 no-person decisions at that rate.
  • Fairness: both checks pass the four-fifths rule. Among claims that could rightly be handled with no person, the worst-served group gets at least 91% of the best-served group's rate.
  • Drift: the certainty score is stable between earlier and later claims (stability index 0.079; under 0.1 is stable).

Risks and open issues

  • Staff time per claim: 74.1 minutes to 74.8 minutes. Watching is a standing cost, at an assumed 10 minutes per audited decision.
  • An audit this size rarely catches a rare error: with 2 wrong decisions among 626, a 10% sample had a 19% chance of seeing any of them, and this draw found 0. This synthetic draw does not validate production accuracy.
  • The test claims carry no protected characteristics, so fairness is checked on channel and customer tenure only. A real deployment must test the groups the law names.
  • Drift is measured inside one synthetic run, where none is expected. This shows the monitor works, not that the model is stable over time.
  • Complaints linked back to decisions are part of this stage and are not built.

Judged by the lab's money rule, oversight only costs. Its value is what it would catch, shown below.

What the audit would catch

2 of the 626 decisions made with no person were wrong. A random audit has to be large, or run for long, to see a rare mistake.

Audit shareDecisions checkedChance of seeing any mistakeDecisions before the accuracy bar is proven
5%3510%15,300
10%7319%7,650
20%13336%3,825

Is it fair?

  • How the claim was submitted: the worst-served group is settled alone at 91% of the best-served group's rate. Passes the four-fifths rule. Channel should not change how a claim is treated.
  • How long they have been a customer: the worst-served group is settled alone at 93% of the best-served group's rate. Passes the four-fifths rule. Tenure feeds the fraud questions, so some gap is possible; a large one would be a finding.
  • Kind of insurance: the worst-served group is settled alone at 69% of the best-served group's rate. Expected to differ: larger claims need more certainty by design. Shown for context, not as a fairness test.

The four-fifths ratio is a screening convention from US employment guidance, not a statistical test. These synthetic claims carry no protected characteristics, so the check can cover only how a claim arrived and how long the customer has been one.

Is the model drifting?

Stability index 0.079 between earlier and later claims: stable. Inside one synthetic run no drift is expected, so this shows the monitor works, not that the model is stable over time.

One claim: Life claim for $47,024 C-0479
Before
Handled with no person, and never looked at again
In this step
May be drawn at random for a person to check

The largest claim this version handled alone.

The lesson. The biggest saving came from the work people still do, not from automating more, and it is the least proven: the minutes each new task takes are assumptions until someone times them. A second model and a question to the customer each added a moving part that did not earn its keep.

What each change adds to run, and what it bought

Every step makes the operation harder to run: another model to maintain, another queue to staff, another supplier to govern. The trade is where business judgment shows. Each one's figures are in the table above.

VersionAdds to the operationWorth it?
W0 · Trust the one big answerA decision model in the processNo, abandoned
W1 · Combine the small answersNothing new: a change to codeKeep the idea
W2 · Learn how much to trust each answerA statistical model that must be refitted on real claimsOnly with W3
W3 · Be stricter as amounts riseOne rule: the certainty required rises with the amountYes
+ A quick confirmation for near-certain claims (now for requests only)A new queue and task, and a risk of approving by habitOnly for requests for more information
+ Grade the fraud responseA new check tier that some real frauds meet firstYes, if the check catches fraud
+ A second vote on near-misses (a negative result)A language-model supplier, with claim text leaving for itNo, on this evidence
+ Ask the customer one questionCustomer contact, and untrusted replies entering the processNo: as W4 it failed its test
+ Watch after the decisionAudit capacity, and fairness and drift monitorsYes: the cost of operating

What to run

Run
W3, a quick confirmation for requests for more information only, the graded fraud check (a claim the light check clears goes to a standard review, never an automatic payment) and the audit: trust the model in proportion to the money at stake, give every approval a person must see a full review, and check afterwards. Replayed on the stored answers, with every claim the light check clears given its standard review: 53.7% settled with no person (paid and asked together), 91.4 min of staff time per claim, net +$31,561 per thousand claims. These figures use the corrected way of choosing the cut-off, as in the version that passed, with the payment limit. The replay assumes the light check misses no fraud; in a simulation with assumed rates it must catch at least 72.7% of the claims that warrant an investigation to be worth having.
Leave out
A quick confirmation for approvals: it pays only if the reviewer catches almost every wrong file, far more than people reach. The second opinion: a second vote on the same facts does not earn its supplier and data cost. Asking the customer, frozen as W4, failed its test on fresh claims. Against the graded fraud check alone, the second opinion adds +4 claims settled alone, +$84 per 1,000 claims, and +1 wrong decision. These comparisons use the first way of choosing the cut-off, which failed its test on fresh claims, with no payment limit.
Before it is real
Before any of this is real: a test on real claims rather than synthetic ones, timed human tasks, and reviewers tested on planted wrong files before the quick confirmation carries a claim, even for requests.

One card per workflow, each in the same order

Why each one exists, the idea, the part of the machinery it changes, its result and the lesson it left.

The lab publishes its exact thresholds so the evidence can be checked. An insurer running this design would keep its operating thresholds confidential, shared with its regulator, and add random audits, because published cut-offs invite claims tailored to pass them.

decision modellanguage modelcodechanged The heavy outline marks the one part a version changes.

The main chain: each one fixes the last one's flaw

The same fields on every card, in the same order. The strip is the machinery every version shares, each part in its role's colour.

W0

Trust the one big answer

Built · misses the bar
Why it exists
The starting point: the policy as first specified.
The idea
Settle a claim alone only when the decision model confidently answers the broad question, "what do you recommend?"
Reads
The decision model's one recommendation, plus the evidence checks
Learned from data
No: hand-set
Certainty needed
The same for every claim
Models
Decision model only: Jev, asked once and replayed

The planted fraud sent to a person right

In development 3.4% settled alone (paid and asked together) · 70.59% right · −$100.0k per 1,000 claims

The exact rule

The engine as built: after the fixed rules, act only if the recommendation is "approve" or "ask for information" at or above the confidence threshold, and the documents, photos, estimate and fraud score all pass.

Lesson. One broad answer is rarely confident, and when it is, it is often wrong. Next: stop asking the broad question.

W1

Combine the small answers

Built · loses money
Why it exists
W0 leaned on one broad answer the decision model is bad at.
The idea
Ignore the recommendation. Code reads the narrow answers (complete? photos match? fraud risk?) and acts only if even the weakest one is strong enough.
Reads
Eight narrow answers
Learned from data
No: hand-set
Certainty needed
The same for every claim
Models
Decision model only: Jev, replayed

The planted fraud sent to a person right

In development 36.6% settled alone (paid and asked together) · 100.00% right · −$76.8k per 1,000 claims

The exact rule

After the fixed rules and the specialist check, the approval confidence is the weakest of the narrow answers; act only if it clears the cut-off chosen for 99.5% right on the training claims.

Lesson. Right almost every time, but far too cautious: one doubtful answer sends a good claim to a person, so it costs more staff time than it saves. Next: weigh the answers instead.

W2

Learn how much to trust each answer

Built · paid a fraud
Why it exists
W1 treated every answer as equally important, and the decision model's probabilities sit 9.6 points from the truth on average on these claims.
The idea
A small statistical model, fitted on past claims, learns a weight for each of the fourteen answers and turns them into one score.
Reads
All fourteen answers
Learned from data
Yes: two small logistic models
Certainty needed
The same for every claim
Models
Decision model only: Jev, replayed

The planted fraud paid automatically wrong

In development 57.8% settled alone (paid and asked together) · 99.83% right · −$39.4k per 1,000 claims

The exact rule

Two logistic models score "paying is right" and "asking for information is right"; the likelier one is proposed, and acted on alone if its score clears the cut-off.

Lesson. Accuracy is not enough: it is right almost every time, but one costly mistake, a planted fraud of $49,101, wipes out every saving. Next: ask for more certainty when more money is at stake.

W3

Stricter as amounts rise

Passed twice
Why it exists
W2 needed the same certainty to pay a small claim or a large one.
The idea
The same learned weights as W2, but above $5,000 the doubt allowed shrinks as the amount grows: at twice that amount, the doubt must be half as large.
Reads
All fourteen answers, and the amount
Learned from data
Yes: the same weights as W2, frozen
Certainty needed
Rises with the amount above $5,000
Models
Decision model only: Jev, replayed; asked afresh in each test on fresh claims

The planted fraud sent to a person right

On fresh claims The first version failed. A corrected version passed, then passed a larger, stricter test. The tests in full.

The exact rule

For a payment, the doubt is one minus the score, times the larger of one and the amount divided by $5,000. It pays alone only if one minus the doubt clears the cut-off: 0.813 in the version that passed, chosen on claims its models had not seen. Asking for information pays nothing, so it is never scaled.

Lesson. Holds the accuracy bar on claims it had never seen; it saves money only with its payment limit. Its known weak spot: a few large planted frauds the decision model misreads.

Tried after W3

Each passed in development and earned a test on fresh claims, except W5, which was only ever an experiment.

W3 +

A fifteenth question

Failed on fresh claims
Why it exists
Most of W3's early mistakes were one kind of claim: damage that may predate the incident, which none of the fourteen answers can see.
The idea
Ask the decision model one more question, "could some of the damage predate the incident?", and let W3 weigh the answer.
Reads
Fifteen answers
Learned from data
Yes: refitted with the new answer
Certainty needed
As W3
Models
Decision model: Jev, one more question per claim

On fresh claims 65.6% settled alone (paid and asked together) · 4 wrong · 99.39% right: below the bar Failed

The exact rule

W3 with the new answer as a fifteenth input: the same regression, risk scaling and cut-off rule.

Lesson. The question caught what it was aimed at, but its answer also made W3 ask other customers for information they did not need, and a wrong request counts against accuracy.

W4

Ask the customer

Failed on fresh claims
Why it exists
Many claims W3 sends to a person are missing one piece of information the customer could give.
The idea
When only the certainty cut-off held a claim back, ask the customer one question, then decide again on the reply.
Reads
W3's answers, then the reply, read again by the decision model
Learned from data
Yes: W3's weights
Certainty needed
As W3; a reply settles at most $20,000
Models
Decision model: Jev. Language model: Claude Haiku 4.5, which writes the question and, in this lab, the customer's reply

On fresh claims 59.9% settled alone (paid and asked together) · 1 wrong · net −$43,069 per 1,000: lost money Failed

The exact rule

W3; a claim held back only by the cut-off, with no other block, gets one question. A reply can never lower suspicion, and pays automatically only up to its cap of $20,000.

Lesson. The step that asks the customer made no wrong payment. The loss was a large planted fraud that plain W3 paid by itself, before any question.

W4 +

A payment limit

Failed on fresh claims
Why it exists
W4's one mistake was a large payment.
The idea
Nothing over $20,000 is paid without a person, whichever step proposes it.
Reads
As W4
Learned from data
Yes: W3's weights
Certainty needed
As W4, and a person for every payment over $20,000
Models
As W4

On fresh claims 52.2% settled alone (paid and asked together) · 1 wrong · net −$3,372 per 1,000: lost money Failed

The exact rule

W4, plus a limit on any automatic payment: above $20,000, a person approves.

Lesson. It stopped the large mistake but sent correct large payments to people, which cost more than it saved. A fault in the run also left some claims unasked. Where to draw that line is a policy choice.

W5

A second opinion

Experiments only
Why it exists
The decision model is unsure of some claims a second reader might settle, and a few large frauds fool it.
The idea
The language model reads the claim beside the decision model's findings, and a claim is settled alone only if both agree.
Reads
The claim and the decision model's findings
Learned from data
No
Certainty needed
Agreement, with a certainty floor
Models
Decision model: Jev. Language model: Claude Haiku 4.5, one review per claim

In development On near-misses: 63.0% settled alone (paid and asked together), 99.52% right. On every payment over $20,000 that W3 proposed, it objected to 30 of the 56 W3 would rightly pay, and let 3 of the 20 wrong ones through.

The exact rule

Near-misses: the language model reviews the claims W3 held back only by its cut-off, and a claim is settled on agreement above a floor. Large payments: it reviews every payment over $20,000 that W3 proposes.

Lesson. A negative result, kept on the record: a second vote on the same facts adds little, because it reads the decision model's findings. Never frozen, and never tested on fresh claims.

In the plan, not built

Two ideas from the original plan. Shown so the numbering makes sense, and because W6 is a decision only the business can make.

Plan W3

Ask twice

In the plan, not built
Why it exists
The decision model sometimes gives a different answer to the same question.
The idea
Ask unstable questions again, reworded, and act only when the answers agree.
Reads
The fourteen answers, twice where unstable
Learned from data
No
Certainty needed
Agreement
Models
Decision model: Jev, asked again on some claims

Status Not built. Steadier decisions, but few more of them, and no help with the large-payment risk.

The exact rule

Planned: re-ask the questions whose answers flip, with reworded instructions, and require agreement before acting alone.

Lesson. Its number went to "stricter as amounts rise".

W6

Policy changes

In the plan, not built
Why it exists
Some claims can only be settled alone if the business changes its rules.
The idea
Automatic denial with a written reason and an appeal path, and a specialist's pre-analysis for complex claims.
Reads
As W3
Learned from data
No
Certainty needed
As W3
Models
Small

Status Not built: business and regulatory decisions, off by default, and the business's to make.

The exact rule

Planned as switches, each shown with its consequences, never silently on.

Lesson. The plan expected a few points more at most.

Tested on fresh claims: it failed, then a corrected version passed

Every other figure on this site is cross-validated on the claims the lab was built with. The real test is claims it has never seen. The lab freezes its deciding step, W3, and runs it once on a thousand fresh claims, under a seed drawn only when the run starts. A result is final for the version tested.

W3, first attempt: failed

In development (cross-validated)On fresh claims (held-out)
Paid with no person568 of 1,000592 of 1,000
Asked for documents, with no person58 of 1,00062 of 1,000
Wrong decisions2 (1 paid, 1 asked)6 (4 paid, 2 asked)
Accuracy (honest range)99.68% (98.84% to 99.91%)99.08% (98.01% to 99.58%)
Wrongly paid$4,211$16,194
Net per 1,000 claims+$5,272−$6,363

Paid and asked for documents apart: recomputed from the stored answers and checked against the recorded result.

  • Held-out synthetic claims: failed. 1000 of 1000 analysed; 0 unresolved. Complete.
  • Stored rule: verdict/1; accuracy >= 0.995, net per 1,000 > 0.
  • Ledger cost/liability: $0.121950 (unknown charges retain estimates).
  • accuracy on handled claims is 99.08% (648 of 654), below 99.50%
  • net per 1,000 claims is $-6,363, not above $0
  • Not proven on unseen claims.
  • Synthetic evidence only; passing this gate is not proof of real-world accuracy and does not promote the configuration.
  • Priced with the costs recorded when the version was frozen, the lab's earlier placeholders, not the benchmark costs used elsewhere on this site: $55 an hour; a standard review 25 minutes, a specialist 90, a fraud investigation 120, quarantine 45. The accuracy result does not depend on costs.

Margins: accuracy 99.08% against the 99.5% bar, missed by 0.42 points; it needed 3 fewer wrong decisions to pass. Net −$6,363 per thousand claims against $0, missed by $6,363; every claim analysed.

Model: the lab asked for ~typesafe/jev-latest; every answer came from typesafe/jev-1.13-20260917 (4,000 calls).

Which version, run and configuration
  • Version ec04e73e-87cd-42a2-beda-2db8079a380c; fit 47e0ed13c660d491be0dd4bcdef66fec67e948d52825839068d5f6e53866f825; fit build 2398f73ba3e64619e70ed5c0134b2ab04c26982c.
  • Held-out run 1f7861a9-6a70-4aeb-a23e-a20a8086160f; slot 86447edd-6344-4eef-8965-d11ead070e7e; result build 2398f73ba3e64619e70ed5c0134b2ab04c26982c.
  • Config c6141e35-affc-46b9-9104-24f1a2ebd04e; config hash 02eaff083d3b91a8404526bc9d899262a7f0a6cd5a10f51894cb744b221238bd; input hash 905140949ebaf573eefdca69717a020a8741fdad87e766e6f795e100ddc90cca.
  • Decision model: asked for ~typesafe/jev-latest; answered by typesafe/jev-1.13-20260917 (4,000 calls).

W3, second attempt: passed

In development (cross-validated)On fresh claims (held-out)
Paid with no person543 of 1,000549 of 1,000
Asked for documents, with no person50 of 1,00044 of 1,000
Wrong decisions1 (1 paid, 0 asked)0 (0 paid, 0 asked)
Accuracy (honest range)99.83% (99.05% to 99.97%)100.00% (99.36% to 100.00%)
Wrongly paid$4,211$0
Net per 1,000 claims+$5,939+$6,790

Paid and asked for documents apart: recomputed from the stored answers and checked against the recorded result.

  • Held-out synthetic claims: passed. 1000 of 1000 analysed; 0 unresolved. Complete.
  • Stored rule: verdict/1; accuracy >= 0.995, net per 1,000 > 0.
  • Ledger cost/liability: $0.122051 (unknown charges retain estimates).
  • Synthetic evidence only; passing this gate is not proof of real-world accuracy and does not promote the configuration.
  • Priced with the costs recorded when the version was frozen, which are the benchmark costs used everywhere else on this site.

Margins: accuracy 100.00% against the 99.5% bar, cleared by 0.50 points; the bar would still have absorbed 2 more wrong decisions. Net +$6,790 per thousand claims against $0, cleared by $6,790; every claim analysed.

Model: the lab asked for ~typesafe/jev-latest; every answer came from typesafe/jev-1.13-20260917 (4,000 calls).

Which version, run and configuration
  • Version 78a03982-77dc-4e32-9c38-e020bd6157e0; fit cd338fa666922d795df0a3c4f1d144f9c7951226672732d3764866ceeb84ee81; fit build f994cb4157d1b4ffd741ee53b868ddf697a58065.
  • Held-out run b8ac5efa-f0fc-41e1-b7a3-ae820582a2ff; slot 32621ed9-b6a5-4701-965b-b74ed276b1d4; result build f994cb4157d1b4ffd741ee53b868ddf697a58065.
  • Config 1a32a6d0-d4c6-4293-ab98-5b70c158d18a; config hash 2f47180ec0a9036ae90889dff5227487dfb6ac37f38ba0f7a3a2a5bfe55074bc; input hash eb2edb650530b7eb089fe811380dc0926c1ddcd71eac77de5bd14d8e68b3c53a.
  • Decision model: asked for ~typesafe/jev-latest; answered by typesafe/jev-1.13-20260917 (4,000 calls).

Then confirmed on a larger set of fresh claims: passed

A version that passes may take one larger confirmation, on two thousand more fresh claims, declared before its seed is drawn. It is judged more strictly: the bottom of the accuracy range, not the headline figure, must clear the bar. A headline figure can clear the bar by luck on a few hundred claims; the bottom of the range clears it only when the claims are many and the mistakes few. A confirmation never changes the test result it follows.

On 2,000 more fresh claims
Paid with no person1,095 of 2,000
Asked for documents, with no person95 of 2,000
Wrong decisions1 (1 paid, 0 asked)
Accuracy (honest range)99.92% (99.53% to 99.99%)
Bottom of the range, judged against the bar99.53% against 99.5%
Wrongly paid$954
Net per 1,000 claims+$12,613

Paid and asked for documents apart: recomputed from the stored answers and checked against the recorded result.

  • Held-out confirmation on synthetic claims: passed. 2000 of 2000 analysed; 0 unresolved. Complete.
  • Stored rule: verdict/2; the lower end of the 95% accuracy range >= 0.995, net per 1,000 > 0.
  • Ledger cost/liability: $0.244032 (unknown charges retain estimates).
  • Synthetic evidence only; a confirmation neither changes the gate result nor promotes the configuration.
  • Priced with the costs recorded when the version was frozen, which are the benchmark costs used everywhere else on this site.

Margins: the bottom of the accuracy range 99.53% against the 99.5% bar, cleared by 0.03 points; one more wrong decision would have failed it. Net +$12,613 per thousand claims against $0, cleared by $12,613; every claim analysed.

Model: the lab asked for ~typesafe/jev-latest; every answer came from typesafe/jev-1.13-20260917 (8,000 calls).

Which confirmation run
  • Confirmation run f559ede7-133b-4688-b202-ea1e13177b9f; slot f4d5cc1a-e7f3-4096-8f4d-4d872a4a25cb; result build 02f965d7726b1d4fb227df250913366ed1ce14e4.
  • Config 1a32a6d0-d4c6-4293-ab98-5b70c158d18a; config hash 2f47180ec0a9036ae90889dff5227487dfb6ac37f38ba0f7a3a2a5bfe55074bc; input hash 9c02a59f8ca998b2da608b17063bab275bde68e728bbf31a1ee989e47dbc48ea.
  • Decision model: asked for ~typesafe/jev-latest; answered by typesafe/jev-1.13-20260917 (8,000 calls).

A test declared in advance, with the payment limit: passed

Declared before it ran, with a public timestamp, for the version as it runs now, payment limit included, on fresh claims drawn after the limit existed; the declaration itself is published with the result. It is judged on the bottom of the accuracy range, and it changes neither the version's result nor its standing.

To check it: the declaration named drand round 32600518, and its timestamp is in Bitcoin block 968849. The declaration, its proof and the round's beacon are in the published record. What this proves: each test's version and settings were declared before its seed could be known. The declaration is in a Bitcoin block at least two hours before the public drand round it names, and the seed is computed from that round. The log is hash-chained and lists every declared test with its status, so a test declared and then dropped still shows. What it can't prove: declarations newer than the latest published one could still be held back, and the block header is checked offline, so confirm its height in any Bitcoin explorer.

On 3,000 fresh claims
Paid with no person1,486 of 3,000
Asked for documents, with no person127 of 3,000
Wrong decisions1 (1 paid, 0 asked)
Accuracy (honest range)99.94% (99.65% to 99.99%)
Bottom of the range, judged against the bar99.65% against 99.5%
Wrongly paid$3,293
Net per 1,000 claims+$4,059

Paid and asked for documents apart: recomputed from the stored answers and checked against the recorded result.

  • Held-out config-check on synthetic claims: passed. 3000 of 3000 analysed; 0 unresolved. Complete.
  • Stored rule: verdict/2; the lower end of the 95% accuracy range >= 0.995, net per 1,000 > 0.
  • Ledger cost/liability: $0.366392 (unknown charges retain estimates).
  • Run under the configuration that decided when it ran, with no automatic payment above $20,000.
  • Synthetic evidence only; a config-check changes neither the version's status nor its promotion.
  • Priced with the costs recorded when the version was frozen, which are the benchmark costs used everywhere else on this site.

Margins: the bottom of the accuracy range 99.65% against the 99.5% bar, cleared by 0.15 points; the bar would still have absorbed 1 more wrong decision. Net +$4,059 per thousand claims against $0, cleared by $4,059; every claim analysed.

Model: the lab asked for ~typesafe/jev-latest; every answer came from typesafe/jev-1.13-20260917 (12,000 calls).

Which run
  • Config-check run fa1c09db-78ac-4527-a9e4-d50c8e2bbc85; slot 78456d61-116c-4a64-a67f-dc039ebe4d8b; result build 34060d27e272276ac3f989cd447521a36e4a84d2.
  • Config 6491718a-ed24-4ae9-b950-3a11d7af7600; config hash 5b7f6a57bf9cacf3799f96a1e4c1c35f0b5307a3996eeaa90b30fb0ee5e7341e; input hash 2b3ec422f4794ea83669a096966f2cfc36f86f5dbbaa5377a6fa1de5ac305627.
  • Decision model: asked for ~typesafe/jev-latest; answered by typesafe/jev-1.13-20260917 (12,000 calls).

What changed between the attempts: only how the certainty cut-off is chosen. The first version picked the lowest cut-off that was right often enough on the claims it had just learned from, and a model is always more sure of claims it has seen. The second version scores those same claims with models fitted without them, and picks its cut-off from those scores. The failure prompted the change, but the new cut-off was chosen on the lab's own claims, never on the held-out ones, and the second version was tested on a new set of fresh claims.

The corrected version made no wrong decisions and a profit, though it settles fewer claims alone. These are synthetic claims: a pass starts the evidence, it does not end it.

Each result is final for its frozen version: the first failure stays on the record. A second attempt needed a new frozen version, a newly declared test slot and a new spending approval. The declarations for the tests so far are kept in the lab's private records. From here on, each test's declaration, including the frozen version's hash, is timestamped publicly before the test runs. The first test declared this way has run: its declaration was timestamped in a Bitcoin block before the public drand beacon round it named, and its fresh claims were drawn from that round.

It then passed a larger, stricter test: the bottom of its range, 99.53%, cleared the 99.5% bar.

The development figures for every version still use the first way of choosing the cut-off, except the recommended design, which has been replayed the corrected way, with the payment limit. The version that passed is stricter, so it settles fewer claims with no person than the development W3 shows.

The designs built after W3

Three designs were frozen after W3 on the same development claims, each with its own test on a thousand fresh claims. All three failed. Their records are kept apart from W3's, and none changes W3's result.

W3 + a fifteenth question Failed

Asks the decision model whether the damage could predate the incident: the one kind of claim that caused most of W3's early mistakes. The question caught what it was aimed at. But its answer also made W3 ask other customers for information they did not need, and a wrong request counts against accuracy.

In development (cross-validated)On fresh claims (held-out)
Settled with no person (paid and asked together)654 of 1,000656 of 1,000
Wrong decisions14
Accuracy (honest range)99.85% (99.14% to 99.97%)99.39% (98.44% to 99.76%)
Wrongly paid$3,790$5,634
Net per 1,000 claims+$10,630+$13,196
  • Held-out synthetic claims: failed. 1000 of 1000 analysed; 0 unresolved. Complete.
  • Stored rule: verdict/1; accuracy >= 0.995, net per 1,000 > 0.
  • Ledger cost/liability: $0.146820 (unknown charges retain estimates).
  • accuracy on handled claims is 99.39% (652 of 656), below 99.50%
  • Not proven on unseen claims.
  • Synthetic evidence only; passing this gate is not proof of real-world accuracy and does not promote the configuration.
  • Priced with the costs recorded when the version was frozen, which are the benchmark costs used everywhere else on this site.

Margins: accuracy 99.39% against the 99.5% bar, missed by 0.11 points; it needed 1 fewer wrong decision to pass. Net +$13,196 per thousand claims against $0, cleared by $13,196; every claim analysed.

Model: the lab asked for ~typesafe/jev-latest; every answer came from typesafe/jev-1.13-20260917 (5,000 calls).

Which version, run and configuration
  • Version 37e208ce-d730-49ce-af47-9e8126076dfc; fit ebcb4ac8ae858cc9f1e96862939d91de8293f406981aecf051a25f13acc9a605; fit build 33b0bcbecfda2584d875732cf6bc4099849ddd36.
  • Held-out run dc4f870c-f86f-4e3a-b3ad-40bd3dd79044; slot afc69a57-644f-4766-8ac7-5ab703c5d0b1; result build 33b0bcbecfda2584d875732cf6bc4099849ddd36.
  • Config 6f853dbf-41dd-4d0d-81a9-59e454097577; config hash 6a2251bd596b3265838ea1ef92ba847ef7e2647443002eabcda5ee1ae030a8b0; input hash 3f6ee38230c81cd226a8cdc81dd832e676a1d16bf9a38181931e54903c55614d.
  • Decision model: asked for ~typesafe/jev-latest; answered by typesafe/jev-1.13-20260917 (5,000 calls).

W4: ask the customer Failed

When one piece is missing, ask the customer for it, then decide again. A reply can settle a claim only up to a cap. The step that asks the customer made no wrong payment, and caught every attempt to manipulate it, against five fixed test phrasings; subtler attacks are not yet tested. The claim that failed the test was one plain W3 paid by itself.

In development (cross-validated)On fresh claims (held-out)
Settled with no person (paid and asked together)614 of 1,000599 of 1,000
Wrong decisions11
Accuracy (honest range)99.84% (99.08% to 99.97%)99.83% (99.06% to 99.97%)
Wrongly paid$4,211$49,999
Net per 1,000 claims+$5,869−$43,069
  • Held-out synthetic claims: failed. 999 of 1000 analysed; 1 unresolved. Incomplete evidence cannot pass.
  • Stored rule: verdict/1; accuracy >= 0.995, net per 1,000 > 0.
  • Ledger cost/liability: $0.429629 (unknown charges retain estimates).
  • net per 1,000 claims is $-43,069, not above $0
  • the evidence is incomplete: not every claim was analysed, and an incomplete run cannot pass
  • Not proven on unseen claims.
  • Synthetic evidence only; passing this gate is not proof of real-world accuracy and does not promote the configuration.
  • Priced with the costs recorded when the version was frozen, which are the benchmark costs used everywhere else on this site.

Margins: accuracy 99.83% against the 99.5% bar, cleared by 0.33 points; the bar would still have absorbed 1 more wrong decision. Net −$43,069 per thousand claims against $0, missed by $43,069; only 999 of 1,000 claims analysed, short of the complete run the rule requires.

Model: the lab asked for ~typesafe/jev-latest; every answer came from typesafe/jev-1.13-20260917 (4,624 calls).

Which version, run and configuration
  • Version 82d30d60-d225-4047-9779-58816c46de28; fit 82640130a404c5e392a208abb355390b46ca147ff5c866a3604551f71b44127c; fit build 78c42acac3cb24c5d544f5444e0504638d65f261.
  • Held-out run 99e4ee1a-9911-469b-95b9-270752bbd986; slot 2ecf46bf-333c-4e11-8b4b-df1062936df8; result build 78c42acac3cb24c5d544f5444e0504638d65f261.
  • Config a9173284-36f4-40ed-b5ec-5862569905fe; config hash dfa3c1d3424ff8742cb9a596d30959535aa95780399eddf0d19df06e98d9c9c1; input hash 7f28c07788aa6fc447d6318d61f7793d0685aaec3f0bc60e5bda64a26508f1b7.
  • Decision model: asked for ~typesafe/jev-latest; answered by typesafe/jev-1.13-20260917 (4,624 calls).

W4 + a payment limit Failed

Nothing over the limit is paid without a person, whichever step proposes it. The limit stopped the large mistake, but sent correct large payments to a person, which cost more staff time than it saved. A fault in the run also left some claims unasked, so its evidence is incomplete.

In development (cross-validated)On fresh claims (held-out)
Settled with no person (paid and asked together)558 of 1,000522 of 1,000
Wrong decisions11
Accuracy (honest range)99.82% (98.99% to 99.97%)99.81% (98.92% to 99.97%)
Wrongly paid$4,211$2,322
Net per 1,000 claims+$1,949−$3,372
  • Held-out synthetic claims: failed. 905 of 1000 analysed; 95 unresolved. Incomplete evidence cannot pass.
  • Stored rule: verdict/1; accuracy >= 0.995, net per 1,000 > 0.
  • Ledger cost/liability: $0.261150 (unknown charges retain estimates).
  • net per 1,000 claims is $-3,372, not above $0
  • the evidence is incomplete: not every claim was analysed, and an incomplete run cannot pass
  • Not proven on unseen claims.
  • Synthetic evidence only; passing this gate is not proof of real-world accuracy and does not promote the configuration.
  • Priced with the costs recorded when the version was frozen, which are the benchmark costs used everywhere else on this site.

Margins: accuracy 99.81% against the 99.5% bar, cleared by 0.31 points; the bar would still have absorbed 1 more wrong decision. Net −$3,372 per thousand claims against $0, missed by $3,372; only 905 of 1,000 claims analysed, short of the complete run the rule requires.

Model: the lab asked for ~typesafe/jev-latest; every answer came from typesafe/jev-1.13-20260917 (4,284 calls).

Which version, run and configuration
  • Version 9beb673c-403a-4afc-818d-2f0546085370; fit afb2340e5fee7d98410a679616c0498d47b4c7d5995b24a19c8ff8712725d35a; fit build 5a0a26c0a5360e2a1e1615ca062c81025b1242da.
  • Held-out run 59468f88-75c3-4ff3-8905-768ac1650f81; slot 392f1438-733f-4419-ae53-b3f3f89f3bbc; result build 5a0a26c0a5360e2a1e1615ca062c81025b1242da.
  • Config 937a2f30-07fc-4dfc-b781-18f323f58199; config hash 8d5b4277323d8583a33ba4a6fc0308c507cad55d6db234c58cd2f3204e608d63; input hash f8dd984ab9866d5c1dd4ed709f46f874a47ffde419cd4614785184ee5c0005bc.
  • Decision model: asked for ~typesafe/jev-latest; answered by typesafe/jev-1.13-20260917 (4,284 calls).

Every set of fresh claims, pooled

The version that passed was also replayed, frozen and never refitted, on the fresh claims drawn for the other tests: the first W3 attempt's and the three later designs'. Its own test and confirmation replay exactly to what they recorded. The version in shadow mode adds one rule: nothing over $20,000 is paid without a person.

With the payment limit, as it runs in shadow mode
Fresh claims fromClaimsPaid aloneAsked for documentsWrongNet per 1,000Range
W3, first attempt1,00049.9%4.4%1+$4,410−$5,250 to +$13,370
Its own test (W3, second attempt)1,00050.7%4.4%0+$3,850−$5,740 to +$13,370
Its confirmation2,00050.7%4.8%1+$9,813+$3,465 to +$15,680
W3 + a fifteenth question1,00048.2%3.6%0+$8,820$0 to +$16,660
W4: ask the customer1,00047.4%4.6%0+$1,680−$8,050 to +$10,500
W4 + a payment limit1,00047.3%4.0%1−$1,692−$12,612 to +$8,120
Its test declared in advance, with the payment limit3,00049.5%4.2%1+$4,059−$1,729 to +$9,519
All sets, pooled10,00049.4%4.3%4+$4,887+$1,861 to +$7,951
Without the limit, as it was tested
Fresh claims fromClaimsPaid aloneAsked for documentsWrongNet per 1,000Range
W3, first attempt1,00055.0%4.4%1+$7,980−$1,470 to +$17,150
Its own test (W3, second attempt)1,00054.9%4.4%0+$6,790−$2,730 to +$15,750
Its confirmation2,00054.8%4.8%1+$12,613+$6,361 to +$18,563
W3 + a fifteenth question1,00054.0%3.6%0+$12,880+$4,340 to +$21,000
W4: ask the customer1,00053.2%4.6%1−$43,699−$147,827 to +$12,390
W4 + a payment limit1,00053.1%4.0%1+$2,368−$8,214 to +$12,518
Its test declared in advance, with the payment limit3,00054.9%4.2%1+$7,792+$2,449 to +$13,300
All sets, pooled10,00054.4%4.3%5+$3,492−$8,139 to +$10,728

Net ranges are the middle ninety-five percent of 2,000 seeded resamples of the claims, at the benchmark staff costs. They reflect how often large planted frauds occur in these sets of fresh claims, not how often they would occur elsewhere.

With no model, on rules alone, the same W3 method settles 4.7% of the development claims with no person and nets +$3,290 per thousand claims; with the decision model, 62.6% and +$8,809.

Accuracy against the share settled, if the line moved

How often W3 is right as its line is moved, on 10,000 pooled fresh claims and, dashed, on the development claims. The version keeps the line it was frozen with; nothing here chooses another.

Accuracy against the share of claims settled with no personSolid: the fresh claims, with the honest range shaded. Dashed: the development claims. The dot is W3's frozen line. The share axis starts at 30%: below it so few claims are settled that their accuracy says little.
93%94%95%96%97%98%99%100%30%40%50%60%70%Share of claims settled with no person (paid and asked together)the accuracy bar, 99.5%W3's frozen line

W3's frozen line, a certainty cut-off of 0.813, settles 53.7% of these fresh claims with no person (paid and asked together), 99.93% right (range 99.81% to 99.97%).

The most W3 could settle with no person while keeping each accuracy target, and at which cut-off: by the measured accuracy, and by the bottom of its range.
Accuracy targetFresh claims, by the measured accuracyFresh claims, by the bottom of the rangeDevelopment claims, by the measured accuracy
99.0%58.5% (cut-off 0.575)57.9% (cut-off 0.626)58.7% (cut-off 0.615)
99.5%57.0% (cut-off 0.689)56.3% (cut-off 0.720)57.8% (cut-off 0.649)
99.9%54.8% (cut-off 0.775)52.9% (cut-off 0.834)53.5% (cut-off 0.780)
  • Descriptive only. This curve is drawn on held-out claims after the fact; a cut-off chosen from it would be fitted to them, so no cut-off is chosen here, and the version keeps the one it was frozen with. The payment limit was itself set after one of these sets (see the retrospective note), so this curve inherits that.
  • Only W3's line moves: every hard rule, and the payment limit, stays exactly as frozen at every point.
  • The development curve is development evidence: the same design cross-validated on the evaluation claims, the claims the designs were built on. Where the two curves part is how far development flattered the design.

The people it takes

Staff time, turned into the number a business plans around: people. Each figure is one full-time person doing claim work.

How many people each version needs

The evidence counts the staff work behind every version, task by task: 122.0 min of staff time per claim by hand, 91.4 min in the recommended version, audits included. Change any assumption below; the work stays what the evidence measured, only its price and the headcount move.

Minutes per task
By hand127 people122.0 min of staff time per claim
Recommended version95 people91.4 min of staff time per claim
Same team, more claims1.3×25% less staff time

Where the people go

Where the people goBy handRecommended
Standard review5217
Quick confirm–0.1
Light fraud check–1.6
Fraud investigation6161
Specialist review1313
Quarantine1.82.0
Audits after the decision–0.8
Claims per yearPeople by handPeople, recommendedSaved a year
10,000139.5$358k
100,00012795$3.6M
1,000,0001,271952$35.8M

Full-time equivalents doing claim work only: supervisors, managers and queue slack would make both columns larger. The recommended version's fraud investigations do not shrink: they are the claims that genuinely need one, and after automation they are most of the remaining work.

Automation moves people to the claims that need them

The recommended design does not remove people from claims. It takes routine claims off their desks: the clean ones are settled with no person, and a request for more information the model is nearly sure of gets a quick confirmation instead of a full review. Every approval a person must see gets a full review.

The same team can take on more claims before it has to grow.

Where the cost assumptions come from

AssumptionUsed hereBenchmarkSource
Standard review60 minutes30 to 90 minutes; simple home claims take 1 to 3 hours On its test on 2,000 fresh claims, W3 still saves money while this takes at least 43.3 min; nothing published yet confirms or rules that out.Industry practice; no published touch-time measurement (weak)
Specialist review6 hours4 to 10 hours or more for complex filesIndustry practice (weak)
Fraud investigation9 hours174 accepted cases per investigator per year: about 9 to 10 hours eachCoalition Against Insurance Fraud, SIU Benchmarking Study 2024
Quarantine (security review)1 hourNo benchmark: set equal to a standard reviewLab assumption
Staff cost$70 per productive hourMedian wage $37.51 an hour plus benefits (30% of pay): about $54 per paid hour, about $70 per productive hourUS Bureau of Labor Statistics, occupation 13-1031 (2025) and Employer Costs for Employee Compensation
Quick confirm, light check, audit3, 20 and 10 minutesNo benchmarkLab assumptions
Wait for a decision by hand5 daysDecision within 21 days of proof of loss (regulation); final payment on home claims about 41 days. Five days is likely optimistic for people, and no source measures the time to a decision on routine claimsNAIC Model Regulation 902; J.D. Power 2026 Property Claims Study (the five days: lab assumption, weak)

What is not proven

A pass on fresh synthetic claims shows the result reproduces on claims from the same generator, not that it carries over to real ones. The recommended design has been replayed with the corrected way of choosing the cut-off; the other steps built around W3 have not. The evidence adds:

  • A frozen W3 has held-out evidence; every other figure here is development evidence on synthetic claims.
  • Staff minutes per task and the $70 cost of a productive hour are benchmark mid-points, and the short lanes' and audits' minutes are the lab's assumptions: none are measurements.
  • A person is assumed to decide every claim they see correctly; people's own errors are not measured.
  • An automated request for information is a next action only: the customer's reply and the claim's final outcome are not recorded.
  • The second opinion rests on real calls for 110 near-miss claims only.
  • Modeled net is staff capacity at assumed costs, not realized cash savings; this is a demonstration, not a customer deployment.

The 99.5% accuracy bar is the lab's starting assumption, set when the lab began and taken from no outside source: it is not an industry standard. Published figures on people's accuracy measure other things, and people's own accuracy on these claims is not measured.

What the research says

Each way the lab tests a design guards against a risk that published research names, and the first W3's failure is what that research predicts. The same research also cuts against two of the steps built around W3, and names the limit that matters most.

Where the method follows it

Choosing the cut-off on claims the models had not seen
Tuning a setting and scoring it on the same claims flatters the score. That is how the first W3 failed, and why the corrected one chooses its cut-off on claims its models never saw. Varma & Simon, 2006, BMC Bioinformatics; Cawley & Talbot, 2010, Journal of Machine Learning Research
Many splits of the claims, not one
One split can be lucky. The spread across many splits is what a design is judged on before it earns a test. Dietterich, 1998, Neural Computation; Bouthillier and others, 2021, MLSys
A frozen design and a single test
Choices made after seeing a result inflate it, and a test set that is used again leaks into the design. So each version is frozen before its seed is drawn and tested once. Nosek and others, 2018, PNAS; Dwork and others, 2015, Science
The honest range around accuracy
The Wilson interval stays reliable near perfect accuracy, where the textbook interval does not. The stricter test judges the bottom of that range. Brown, Cai & DasGupta, 2001, Statistical Science
Money as well as accuracy
Counting mistakes alone misleads when mistakes differ in cost: W2 is right almost every time and still loses money. Elkan, 2001, IJCAI
W3: certainty scaled to the money at stake
A textbook decision with the option to refer: act alone only when the expected cost of being wrong is small, and otherwise hand the claim to a person. Chow, 1970, IEEE Transactions on Information Theory; Correa Bahnsen, Aouada & Ottersten, 2015, Expert Systems with Applications

Where it cuts against a step

The quick confirmation
A review of more than a hundred experimental studies found that, on decision tasks, a person working with an AI did worse on average than the better of the two alone, most of all when the automated system alone was the stronger. A person approving a prepared file has that shape. Its saving stays a hypothesis until reviewers are tested with planted wrong files. Vaccaro, Almaatouq & Malone, 2024, Nature Human Behaviour
The second opinion
Combining two deciders pays when they are diverse, when their mistakes differ. When two language models both err, they often give the same wrong answer, and more accurate models more so, and this one also reads the decision model's findings. A second vote on the same facts is weak evidence: the certainty floor does the work, and the lab keeps it as a negative result. A language model is worth its cost where it reads or checks something the decision model cannot. Steyvers, Tejeda, Kerrigan & Smyth, 2022, PNAS; Kim, Garg, Peng & Garg, 2025, ICML

The limit that matters most

Reproduced, not yet carried over
The tests on fresh claims show that the result reproduces on new claims from the same generator. Whether it carries over to real claims is a separate question, and reporting guidance expects both to be stated. Nothing here tests the second yet. Justice, Covinsky & Berlin, 1999, Annals of Internal Medicine; Collins and others, 2024, TRIPOD+AI, BMJ

What was tested, and how

1,000 synthetic claims, answered by ~typesafe/jev-latest through jev-openrouter. 5-fold cross-validation: every claim is scored by a workflow whose cut-offs and weights were fitted on the other four folds.

The claims are generated by code from a fixed seed, across auto, home, life and commercial insurance; the generator sets each claim's right answer and plants the frauds, manipulation attempts and missing papers.

Disclosure: the lab has no relationship with TypeSafe, the company that makes Jev. Jev was used through OpenRouter, and the lab's owner paid for every model call, Jev's and the language models', personally. The lab's owner designed, ran and judged every test, working with AI coding assistants; nothing here has yet been replicated independently.

Evidence: run 6827826e-b962-4bd8-bdb3-b163712fd660, exported from build 7c88683e4bce, dataset 96f9d5e30b51. Synthetic claims and stated assumptions: a demonstration, not a customer deployment.

Costs are benchmark mid-points, stated once

Staff minutes per task and the cost of a productive hour are mid-points of published industry benchmarks, listed below with their sources. Four have no benchmark and are the lab's own assumptions. None are any insurer's records, so read the dollar figures for direction and relative size, not as a savings forecast.

How each number is known, strongest first

  • Held-out testA frozen version run once on fresh claims it never saw, under a seed drawn when the run started. The strongest evidence here, and the only kind that can pass or fail the design.W3, 2 frozen versions
  • Tested with real callsA real second model was called on the claims it would see.+2nd
  • Measured in the engineThe rules as first built, run once over every claim. Measured differently from the later steps, so part of the jump to W1 is the method.W0
  • Cross-validatedEach claim is graded by settings tuned on the other claims, never on itself.W1, W2, W3
  • Replayed, assumed minutesThe stored answers routed through the new process for real; the minutes each new task takes are assumed.+QC, +Fraud, +Audit
  • Simulated repliesCustomer replies written by a language model, including deliberately misleading ones.+Ask
  • Assumed starting pointThe by-hand process, priced with the lab's assumed minutes. Nothing is measured.By hand

The evidence behind every figure on this site is published as it was exported, byte for byte: the manifest names each file and its hash.