How the numbers were made, and how far to trust each one.
Figures last updated , from the published evidence. What changed
1How the lab runs
From synthetic claims to a published result, in eleven steps
The decision model is asked once and its answers are stored, so every workflow can be tried again for free. Only the final test asks it afresh, once, on claims no one has seen. Swapping either model is one setting, from a reviewed list, and every model goes through the same steps.
Once, paid
1Synthetic claimsA seeded generator: 1,000 claims with known right answers; frauds, manipulation and missing papers planted
2The decision model answersJev, through OpenRouter's decisions endpoint, in TypeSafe's System One format. It could as well be TypeSafe direct, a free practice model, or a language model wrapped to answer in the same format
3Every paid call approvedA ledger approves each paid run in advance, within a fixed budget for the whole lab
4Answers storedEvery answer and probability kept, with the model and question versions
Free, as often as needed
5Replay every workflowEvery workflow reads the same stored answers: only the decision changes
6Cross-validationEach claim is scored by settings tuned on the other claims, never on itself
7AdmissionA design earns a real test only if it passes on at least nine in ten of a hundred different splits of the claims
8FreezeThe weights and the cut-off are written down and never changed again
One shot, paid
9Test on fresh claims1,000 new claims from a seed drawn when the run starts; the decision model asked afresh; the result final
10ConfirmationFor a version that passed: 2,000 more, judged on the bottom of the accuracy range
11Published evidenceExported byte for byte with its hashes; every figure on this site traces to it
The lever. Accuracy and speed move with the model and with the process around it. A new model goes through all eleven steps. A new question to the decision model, or a new step such as asking the customer, goes through the same tests as every workflow before it: nothing reaches this site until it has faced fresh claims.
The decision model, measured on this run
Every call the run made is recorded, with its time and its cost. Its yes-or-no answers were checked against the known right answers: the chart puts the probability it gave beside how often the answer was yes.
Calls
13,278 on 1,100 claims, every test on this run
Time a call
276 ms median; one in twenty over 515 ms
Cost
$0.40 in all: $0.03 per thousand calls
Calibration
9.6 points off on average, over 7,000 yes-or-no answers
Each dot is a tenth of its yes-or-no answers, grouped by the probability it gave: across, the average probability; up, how often the answer was yes. On the diagonal the two agree. At the top the answer was yes more often than it said, and at the bottom less often: mostly more cautious than it needed to be.The ten groups
Answers
What it said, on average
How often yes
700
3.1%
0.0%
700
5.2%
0.6%
700
7.5%
0.6%
700
16.9%
4.1%
700
35.7%
30.1%
700
57.9%
27.0%
700
77.5%
84.3%
700
83.3%
93.6%
700
87.6%
97.6%
700
95.3%
100.0%
2What the claims are made of
What the synthetic claims contain, and how a decision is judged right
Every figure on this site comes from synthetic claims the lab generated. Every set of a thousand is built to the same quotas, so the fresh sets used in the tests can be set beside the claims the designs were developed on. Categories are counted here; how a fraud or a manipulation attempt is built is not described.
What each set of claims is made of, as shares of its claims
Share of claims
Development claims (1,000)
Fresh claims, pooled (10,000)
By line of insurance
Auto
40.0%
40.0%
Home
30.0%
30.0%
Commercial
20.0%
20.0%
Life
10.0%
10.0%
By amount claimed
Under $20,000
75.5%
75.4%
$20,000 to $50,000
16.5%
16.6%
Over $50,000
8.0%
8.0%
Planted by the generator
Fraud to investigate
10.0%
10.0%
Manipulation attempts
5.0%
5.0%
Ambiguous
15.0%
15.0%
Missing documents
5.1%
4.5%
Who must see it
Must reach a person, by rule
21.6%
21.2%
The most any workflow could settle with no person
78.4%
78.8%
The planted shares are the lab's own quotas, not any insurer's mix of claims. Accuracy on the claims settled with no person depends on that mix, and would differ with another.
How a decision is judged right
Every synthetic claim carries an answer key written by the generator from the facts it planted: the dispositions it accepts. A decision is judged against that key, never against a person's or a model's opinion.
Paid with no person
Paying without a person is right when the claim's answer key accepts payment.
Asked for more, with no person
Asking the customer for more is right when the answer key accepts a request for more information, as it does for a claim missing documents.
Sent to a person
Sending the claim to a person is right when the answer key calls for review, investigation or quarantine. An ambiguous claim's key always accepts a person.
Wrong
A settled claim is wrong when its disposition is not one its answer key accepts. Claims sent to a person are not settled and are not counted as wrong.
The routing rule. A claim over $50,000, a claim the answer key marks as fraud to investigate, and a manipulation attempt to quarantine must always reach a person. Their share caps what any workflow can settle with no person.
How the claims are generated
Quotas are exact per 1,000 claims: by line, by category (clean, ambiguous, fraud, manipulation) and the share over $50,000. The seed changes which claims are generated, their amounts, documents and wording, never the quotas. Generator version v1.
The evaluation corpus is generated from one fixed, public seed; every design is developed and cross-validated on it, so its figures are development evidence.
A held-out test's seed does not exist until its run starts, from a slot declared in advance for one frozen version. The tests run before 28 September drew it from the operating system's random source at that moment; from the declaration log v2 on, it is computed from a public drand round the declaration named a day ahead.
A confirmation adds a second 1,000 claims from a seed derived from the first. Held-out corpora carry no hand-written fixtures and no rewordings, and no held-out seed is reused in any pool.
A test declared in advance for the version as it runs now, payment limit included, runs once on 3,000 claims: from its slot's seed and two seeds derived from it, each checked against every earlier seed and registered to that slot.
3The process, step by step
The process, today and new
An adjuster reads every claim, checks the policy by hand and decides. It is careful and slow: most of the waiting is a claim sitting in a queue.
Of 6 steps, people do 5, code 1, the decision model none.
Code
Decision model
People
PeopleDecision modelCodeWhere claims end up
1Code: Claim arrives
2People: Waits in a queue
3People: Adjuster reads everything
4People: Adjuster checks the policy by hand
5People: Adjuster decides
6People: Payment is set up
Where the 1,000 claims end up:
Where claims end up
Step 6, Payment is set up → All 1,000 decided by a person, at step 6
What each step does
Claim arrives
Waits in a queue Most of the 5 days before a decision (an assumed wait) is spent here.
Adjuster reads everything
Adjuster checks the policy by hand
Adjuster decides
Payment is set up
By hand, a claim takes about 122.0 min of staff time and waits about 5.0 days for a decision, at benchmark staff times and the lab's assumed wait.
PeopleDecision modelCodeWhere claims end up
1Code: Claim arrives
2Decision model: Jev answers narrow questions
3Code, a fixed rule: Over $50,000?
4Code, a fixed rule: Manipulation attempt?
5Code, a fixed rule: Fraud signs: how strong?
6Code, a fixed rule: Policy in force?
7Code: Certain enough for the amount?
8Code: Over $20,000?
9People: Quick confirm, requests only
10Code: Pay, or ask for what is missing
11People: Audit, afterwards
Where claims go:
Where claims end up
Step 3, Over $50,000? → a person, with the analysis attached, leaving at step 3
Step 4, Manipulation attempt? → quarantine, leaving at step 4
Step 5, Fraud signs: how strong? → a light check or an investigation, leaving at step 5
Step 6, Policy in force? → a person, who confirms any denial, leaving at step 6
Step 7, Certain enough for the amount? → a standard review by a person, leaving at step 7
Step 8, Over $20,000? → a person approves the payment, leaving at step 8
Step 10, Pay, or ask for what is missing → paid, or asked for documents, with no person, at step 10
What each step does
Claim arrives Code notes the amount and looks up the policy once. No model has been asked anything yet.
Jev answers narrow questions Fourteen typed questions about completeness, evidence and fraud, each answered with a probability. Jev answers; it never decides.
Over $50,000? Fixed in code. No setting and no answer from the decision model can change it.
Manipulation attempt? Code acts on Jev's answer about hidden instructions in the text.
Fraud signs: how strong? Weak signs get a light check by a person first, and a claim it clears gets a standard review. Strong signs, or a failed check, go to a full investigation.
Policy in force? One lookup in the policy system. A person confirms every denial; none is automatic.
Certain enough for the amount? W3 weighs Jev's answers. Above $5,000 the doubt allowed shrinks as the amount grows, and it settles a claim alone only if one minus the doubt clears 0.813.
Over $20,000? Nothing over the limit is paid automatically, whichever step proposes it: a person approves it.
Quick confirm, requests only A request for more information that code is nearly sure of gets one quick look by a person. Every approval a person sees gets a full review.
Pay, or ask for what is missing In about a second, with every step recorded.
Audit, afterwards A random sample of the claims settled with no person is checked by a person afterwards.
The fixed rules, as the engine records them
Claims over $50,000 require mandatory human review
Suspected manipulation of automated processing quarantines the claim
Claims with fraud indicators require investigation
Claims on a policy that is not in force go to a human (denials are never automatic)
4Every version, measured
Every version on one table
Development figures on the 1,000 claims the lab was built with. Rows marked + keep W3's decision and change what happens around it. Settled alone counts paid and asked for documents together: the evidence does not yet show them apart for development. Each version's detail follows. Every row here uses the first way of choosing the cut-off, which failed its test on fresh claims, with no payment limit; the recommended design below uses the corrected way of choosing the cut-off, as in the version that passed, with the payment limit.
Claims settled with no person, paid and asked togethershare of all claimsStaff minutes per claimlower is betterDays until the claimant has a decisionlower is betterNet per 1,000 claims against doing it all by handat benchmark staff costs
The figures behind the charts
Version
Claims settled with no person, paid and asked together
Staff minutes per claim
Days until the claimant has a decision
Net per 1,000 claims against doing it all by hand
Hand Today, by hand
0.0%
122.0 min
5.0 days
$0
W0 Trust the one big answer
3.4%
207.8 min
4.8 days
−$100,030
W1 Combine the small answers
36.6%
187.9 min
3.2 days
−$76,790
W2 Learn how much to trust each answer
57.8%
113.8 min
2.1 days
−$39,441
W3 Be stricter as amounts rise
62.6%
110.9 min
1.9 days
+$8,809
+QC A quick confirmation for near-certain claims (now for requests only)
62.6%
102.2 min
1.4 days
+$18,931
+Fraud Grade the fraud response
62.6%
74.1 min
1.4 days
+$51,761
+2nd A second vote on near-misses (a negative result)
63.0%
74.0 min
1.4 days
+$51,845
+Ask Ask the customer one question
66.3%
72.9 min
1.6 days
+$3,649
+Audit Watch after the decision
62.6%
74.8 min
1.4 days
+$50,909
W0 to W3: turning the answers into a decision
All four change one box in the process: is the evidence strong enough to act alone? The same answers from the same model are read four different ways.
W0Trust the one big answerMeasured in the engineLab verdict: Abandoned: not accurate enough and loses money
The engine acts on its own only when the model confidently answers one broad question, "what do you recommend?", and every hand-set guard agrees.
What changed: The engine as shipped, hand-set 0.8 confidence threshold, every claim, no cross-validation. The model's answers now reach the engine. The engine, its rules and its hand-set thresholds are exactly as first specified.
Settled with no person (paid and asked together)
3.4%
Wrong decisions
10 of 34 settled alone
Staff time per claim
207.8 min
Days to a decision
4.8 days
Wrongly paid
$0
Net per 1,000 claims
−$100,030 against doing it all by hand
Settled with no person 34
Full review 545
Specialist 244
Fraud investigation 145
Quarantine 32
What improved
Handled with no person: 0.0% to 3.4%.
Average wait for a decision: 5.0 to 4.8 days.
Risks and open issues
Staff time per claim: 122.0 minutes to 207.8 minutes.
All 10 errors are valid claims wrongly asked for more information. Accuracy on handled claims is 70.6% (24 of 34), below the 99.5% bar, with $0 wrongly paid. That is why the first design was abandoned.
Almost every claim still reaches a person, so the model adds cost without taking work away.
Sent to a human adjuster, because the model was not confident enough in its one broad answer
Whatever happens to the simplest claim shows what this version can and cannot do.
W1Combine the small answersCross-validatedLab verdict: Improved, does not pass: accurate but loses money
Ignore the broad recommendation. Code combines the model's narrow answers (complete? photos match? fraud risk?) and acts on the weakest one.
What changed: How answers are combined. No new questions, no new model calls.
Settled with no person (paid and asked together)
36.6%
Wrong decisions
0 of 366 settled alone
Staff time per claim
187.9 min
Days to a decision
3.2 days
Wrongly paid
$0
Net per 1,000 claims
−$76,790 against doing it all by hand
Settled with no person 366
Full review 213
Specialist 244
Fraud investigation 145
Quarantine 32
What improved
Handled with no person: 3.4% to 36.6%. The two were measured differently: chapter 1 is the engine as shipped, run once over every claim; this chapter is the lab's cross-validated score over the same answers, so part of the jump is the method, not only the design.
Staff time per claim: 207.8 minutes to 187.9 minutes.
Average wait for a decision: 4.8 to 3.2 days.
Risks and open issues
Accuracy measured at 100.00%, but on 366 handled claims the honest range is 99.0% to 100.0%, so 99.5% is met but not yet proven.
One under-confident answer blocks a good claim; one over-confident answer could pass a bad one.
Customer asked for the missing information at once, with no person (this was the wrong call)
In this step
Sent to a human adjuster, because it was not certain enough for the amount at stake
The previous version got this claim wrong. This one does not.
W2Learn how much to trust each answerCross-validatedLab verdict: Improved, does not pass: accurate but loses money
The model's probabilities do not match how often it is right (mostly too cautious), and some answers matter more than others. A small statistical model, fitted on past claims, learns how to weigh them.
What changed: Replaces hand-set combination rules with 14 learned weights. Still code: no free text, and the hard rules run first.
Settled with no person (paid and asked together)
57.8%
Wrong decisions
1 of 578 settled alone
Staff time per claim
113.8 min
Days to a decision
2.1 days
Wrongly paid
$49,101
Net per 1,000 claims
−$39,441 against doing it all by hand
Settled with no person 578
Full review 245
Fraud investigation 145
Quarantine 32
What improved
Handled with no person: 36.6% to 57.8%.
Staff time per claim: 187.9 minutes to 113.8 minutes.
Average wait for a decision: 3.2 to 2.1 days.
Risks and open issues
Wrongly paid out: $0 to $49,101 (1 wrong decision in 578 handled).
Accuracy measured at 99.83%, but on 578 handled claims the honest range is 99.0% to 100.0%, so 99.5% is met but not yet proven.
Weights are learned from synthetic claims and would need re-fitting on real ones.
Sent to a specialist, because it was flagged as needing a specialist
In this step
Paid in about a second, with no person
The previous version sent this claim to a person. This one handles it correctly.
W3Be stricter as amounts riseCross-validatedLab verdict: Candidate
Same learned weights as W2, but a large claim needs much more certainty than a small one before it is paid without a person.
What changed: The confidence needed now rises with the claim amount. No new model calls.
Settled with no person (paid and asked together)
62.6%
Wrong decisions
2 of 626 settled alone
Staff time per claim
110.9 min
Days to a decision
1.9 days
Wrongly paid
$4,211
Net per 1,000 claims
+$8,809 against doing it all by hand
Settled with no person 626
Full review 197
Fraud investigation 145
Quarantine 32
What improved
Handled with no person: 57.8% to 62.6%.
Staff time per claim: 113.8 minutes to 110.9 minutes.
Average wait for a decision: 2.1 to 1.9 days.
Wrongly paid out: $49,101 to $4,211 (2 wrong decisions in 626 handled).
Risks and open issues
Accuracy measured at 99.68%, but on 626 handled claims the honest range is 98.8% to 99.9%, so 99.5% is met but not yet proven.
Large honest claims wait for a person more often. That is the intended trade.
This step was frozen and tested on fresh claims twice. The first version failed the accuracy bar; a second, with a stricter cut-off, passed. It then passed a larger, stricter confirmation. The figures here use the first version's cut-off. The held-out tests.
Paid in about a second, with no person (this was the wrong call)
In this step
Sent to a human adjuster, because it was not certain enough for the amount at stake
The previous version got this claim wrong. This one does not.
The lesson. Accuracy is the wrong single measure. A version can be right almost every time and still lose money, because one large wrong payment outweighs many small savings. What matters is how likely a mistake is, times how much it costs.
Around W3: the work people still do, and what else was tried
These keep W3's decision and change what happens around it: a short confirmation instead of a full review, a light check before a fraud investigation, a language model's second opinion, a question to the customer, and an audit afterwards.
+A quick confirmation for near-certain claims (now for requests only)Replayed, assumed minutesLab verdict: Kept
A claim the decision model has fully prepared, but is not quite sure enough to pay, does not need a full review. It needs one look at the one doubtful item.
What changed: Adds a lane between "no person" and "full review". It does not raise the no-person rate; it cuts the minutes spent on claims that still reach someone.
Settled with no person (paid and asked together)
62.6%
Wrong decisions
2 of 626 settled alone
Staff time per claim
102.2 min
Days to a decision
1.4 days
Wrongly paid
$4,211
Net per 1,000 claims
+$18,931 against doing it all by hand
Settled with no person 626
Quick confirmation 168
Full review 29
Fraud investigation 145
Quarantine 32
What improved
Staff time per claim: 110.9 minutes to 102.2 minutes.
Average wait for a decision: 1.9 to 1.4 days.
The no-person rate, accuracy and wrongly-paid figures are unchanged by design: this chapter only changes what happens to claims that still reach someone.
Risks and open issues
60 of 168 quick confirms were overridden because the prepared suggestion was not right for the claim. Each one pays for the quick look and the full work.
A quick confirm is assumed to take 3 minutes and wait 1 day. Both are placeholders, not measurements.
The replay assumes the person always catches a wrong suggestion. People may approve by habit when the file looks finished.
Research on people working with automated decisions predicts approval by habit: on decision tasks, a person and an AI together did worse on average than the better of the two alone, and people miss a wrong answer most when it is rare. The replay cannot show whether real reviewers would catch a wrong prepared file. Testing reviewers with planted wrong files is a requirement of this lane, not an option.
In a simulated reviewer, with assumed costs on the development claims, a quick confirmation beats a full review only if the reviewer catches at least 99.6% of wrong files: 51 of the 110 prepared files are wrong, 38 of them approvals worth $977,936. People reviewing automated work miss far more than that. For requests for more information, where a missed wrong file costs only a delay, it pays at any catch rate: 28 requests, 13 of them wrong. So the recommended design keeps the quick confirmation for requests only.
Sent to a human adjuster, because it was not certain enough for the amount at stake
In this step
A person confirms the prepared file in about 3 minutes (assumed)
This claim could have been handled automatically and missed the bar narrowly.
+Grade the fraud responseReplayed, assumed minutesLab verdict: Kept
Today one yes-or-no test sends a claim to a full investigation. Weak signals deserve a light check first, with the evidence file assembled by the decision model either way.
What changed: Replaces the single fraud exit with two levels, and hands investigators a prepared file.
Settled with no person (paid and asked together)
62.6%
Wrong decisions
2 of 626 settled alone
Staff time per claim
74.1 min
Days to a decision
1.4 days
Wrongly paid
$4,211
Net per 1,000 claims
+$51,761 against doing it all by hand
Settled with no person 626
Quick confirmation 168
Light fraud check 78
Full review 29
Fraud investigation 67
Quarantine 32
What improved
Staff time per claim: 102.2 minutes to 74.1 minutes.
The no-person rate, accuracy and wrongly-paid figures are unchanged by design: this chapter only changes what happens to claims that still reach someone.
Risks and open issues
30 claims where an investigation was warranted met the light check first. The replay assumes that check catches them; if it does not, this is where fraud gets through.
A light check is assumed to take 20 minutes. It is a placeholder, not a measurement.
No saving is claimed for handing investigators a prepared file; that would need real timings.
This step models the light check as perfect: every claim that warrants an investigation goes on to one, and every needless referral is cleared by the check alone. In a simulation with assumed rates on the development claims, giving every cleared claim its standard review, the graded check would be worth +$28,980 per thousand claims over no graded check if it missed nothing. It pays only if it catches at least 72.7% of the 23 claims among its 78 that warrant an investigation, or if the standard review after it catches what it misses, in which case it pays at any catch rate. Catching 50% with no backstop, it comes to −$24,079 per thousand claims. The recommended design gives every cleared claim its standard review, never an automatic decision, and its figures count it.
Sent to fraud investigation, because fraud indicators were above the limit
In this step
A light verification clears it and it rejoins the normal flow
An honest claim that tripped a hand-set fraud limit.
+A second vote on near-misses (a negative result)Tested with real callsLab verdict: Kept
Jev is fast and cheap but cautious. For the few claims just under the bar, a language model reads the claim and seven of Jev's findings; agreement, above a certainty floor fitted on other claims, lifts the claim over.
What changed: Adds one model call, for near-misses only, so the cost stays small.
Settled with no person (paid and asked together)
63.0%
Wrong decisions
3 of 630 settled alone
Staff time per claim
74.0 min
Days to a decision
1.4 days
Wrongly paid
$4,211
Net per 1,000 claims
+$51,845 against doing it all by hand
Settled with no person 630
Quick confirmation 164
Light fraud check 78
Full review 29
Fraud investigation 67
Quarantine 32
What improved
Handled with no person: 62.6% to 63.0%.
Staff time per claim: 74.1 minutes to 74.0 minutes.
Tested with real calls to a second model (anthropic/claude-haiku-4.5) on 110 claims, for $0.09 in total.
Risks and open issues
Accuracy measured at 99.52%, but on 630 handled claims the honest range is 98.6% to 99.8%, so 99.5% is met but not yet proven.
Agreement alone is unsafe. With no certainty floor, the second model agreed on 13 claims and 6 were wrong, paying out $148,084. The floor, fitted on other claims, is what keeps those with a person.
1 agreed claim was wrong, paying out $0: two models can be wrong together. Agreement is evidence, not proof.
Only 4 claims were added, so this chapter's own error rate is known very loosely.
Adds a second supplier, and sends claim text to it, for every claim it touches.
The lab's rule kept this step because it cleared the accuracy bar and did not lose money. The rule weighs accuracy and money only. It does not count a second supplier, or claim text leaving the company, and on those this step does not earn its place.
Sent to a human adjuster, because it was not certain enough for the amount at stake
In this step
A second model reads the claim and Jev's findings; if both agree and the claim clears the fitted floor, it is handled
Safe to handle, but Jev alone was not sure enough.
+Ask the customer one questionSimulated repliesLab verdict: Rejected: not accurate enough
Missing information is the largest fixable reason a claim cannot be decided. One specific question, answered by the customer, can turn an unclear claim into a clear one.
What changed: The only proposal that can raise the ceiling on the no-person rate, because it changes what is known about the claim.
Settled with no person (paid and asked together)
66.3%
Wrong decisions
4 of 663 settled alone
Staff time per claim
72.9 min
Days to a decision
1.6 days
Wrongly paid
$53,674
Net per 1,000 claims
+$3,649 against doing it all by hand
Settled with no person 663
Quick confirmation 102
Light fraud check 90
Full review 29
Fraud investigation 67
Quarantine 49
What improved
Handled with no person: 63.0% to 66.3%.
Staff time per claim: 74.0 minutes to 72.9 minutes.
All 17 manipulation attempts hidden in replies were caught and quarantined.
Customer replies were synthesised by a language model (anthropic/claude-haiku-4.5), by design 66 helpful, 26 unhelpful and 17 attempts to manipulate the system. Jev then re-analysed each enlarged claim with real calls. $0.21 in total.
Risks and open issues
Average wait for a decision: 1.4 to 1.6 days.
Wrongly paid out: $4,211 to $53,674 (4 wrong decisions in 663 handled).
Accuracy on handled claims is 99.40%, below the 99.5% bar.
1 planted fraud with a reassuring reply was paid: $49,463. With the helpful-sounding reply added, the claim cleared the certainty bar. Nothing in this flow checks that what a reply describes is real.
1 claim was handled wrongly after the reply, paying out $49,463. A reply is new untrusted text; planted frauds and manipulation attempts were left in the test on purpose.
Code allows one question only: a claim that would need a second question goes to a person.
The replies are invented. This shows what the flow can do with a given reply, not how real customers answer, or whether they answer at all.
A reply is assumed to take 2 days. Claims that still reach a person wait for the reply and then for the queue.
Every manipulation attempt was quarantined, in the claims and in the replies, but all were built from five fixed test phrasings. A defence measured on fixed attacks can look stronger than it is; subtler attacks are not yet tested.
Sent to a human adjuster, because it was not certain enough for the amount at stake
In this step
Paid $49,463 incorrectly after a reassuring synthesised reply.
More convincing text is not verified evidence. This planted fraud crossed the payment threshold after a reassuring synthesised reply; the flow did not verify that the evidence described in the reply was real.
+Watch after the decisionReplayed, assumed minutesLab verdict: Kept
Random human audits can gather evidence about automated decisions; they do not automatically prove production accuracy. This demonstration examines sampling, fairness and drift within a synthetic run.
What changed: Adds a monitoring design to evaluate further. Production validation needs representative data, reliable human checks and evidence over time; complaint tracking is not built. It watches chapter 6, the recommended version: the second opinion (chapter 7) is left out by design, and chapter 8 is not built on.
Settled with no person (paid and asked together)
62.6%
Wrong decisions
2 of 626 settled alone
Staff time per claim
74.8 min
Days to a decision
1.4 days
Wrongly paid
$4,211
Net per 1,000 claims
+$50,909 against doing it all by hand
Settled with no person 626
Quick confirmation 168
Light fraud check 78
Full review 29
Fraud investigation 67
Quarantine 32
What improved
A 10% random audit checks 73 no-person decisions out of 626. Once 765 audited decisions in a row are clean, "at least 99.5% accurate" is supported by evidence: about 7,650 no-person decisions at that rate.
Fairness: both checks pass the four-fifths rule. Among claims that could rightly be handled with no person, the worst-served group gets at least 91% of the best-served group's rate.
Drift: the certainty score is stable between earlier and later claims (stability index 0.079; under 0.1 is stable).
Risks and open issues
Staff time per claim: 74.1 minutes to 74.8 minutes. Watching is a standing cost, at an assumed 10 minutes per audited decision.
An audit this size rarely catches a rare error: with 2 wrong decisions among 626, a 10% sample had a 19% chance of seeing any of them, and this draw found 0. This synthetic draw does not validate production accuracy.
The test claims carry no protected characteristics, so fairness is checked on channel and customer tenure only. A real deployment must test the groups the law names.
Drift is measured inside one synthetic run, where none is expected. This shows the monitor works, not that the model is stable over time.
Complaints linked back to decisions are part of this stage and are not built.
Judged by the lab's money rule, oversight only costs. Its value is what it would catch, shown below.
What the audit would catch
2 of the 626 decisions made with no person were wrong. A random audit has to be large, or run for long, to see a rare mistake.
Audit share
Decisions checked
Chance of seeing any mistake
Decisions before the accuracy bar is proven
5%
35
10%
15,300
10%
73
19%
7,650
20%
133
36%
3,825
Is it fair?
How the claim was submitted: the worst-served group is settled alone at 91% of the best-served group's rate. Passes the four-fifths rule. Channel should not change how a claim is treated.
How long they have been a customer: the worst-served group is settled alone at 93% of the best-served group's rate. Passes the four-fifths rule. Tenure feeds the fraud questions, so some gap is possible; a large one would be a finding.
Kind of insurance: the worst-served group is settled alone at 69% of the best-served group's rate. Expected to differ: larger claims need more certainty by design. Shown for context, not as a fairness test.
The four-fifths ratio is a screening convention from US employment guidance, not a statistical test. These synthetic claims carry no protected characteristics, so the check can cover only how a claim arrived and how long the customer has been one.
Is the model drifting?
Stability index 0.079 between earlier and later claims: stable. Inside one synthetic run no drift is expected, so this shows the monitor works, not that the model is stable over time.
The lesson. The biggest saving came from the work people still do, not from automating more, and it is the least proven: the minutes each new task takes are assumptions until someone times them. A second model and a question to the customer each added a moving part that did not earn its keep.
What each change adds to run, and what it bought
Every step makes the operation harder to run: another model to maintain, another queue to staff, another supplier to govern. The trade is where business judgment shows. Each one's figures are in the table above.
W3, a quick confirmation for requests for more information only, the graded fraud check (a claim the light check clears goes to a standard review, never an automatic payment) and the audit: trust the model in proportion to the money at stake, give every approval a person must see a full review, and check afterwards. Replayed on the stored answers, with every claim the light check clears given its standard review: 53.7% settled with no person (paid and asked together), 91.4 min of staff time per claim, net +$31,561 per thousand claims. These figures use the corrected way of choosing the cut-off, as in the version that passed, with the payment limit. The replay assumes the light check misses no fraud; in a simulation with assumed rates it must catch at least 72.7% of the claims that warrant an investigation to be worth having.
Leave out
A quick confirmation for approvals: it pays only if the reviewer catches almost every wrong file, far more than people reach. The second opinion: a second vote on the same facts does not earn its supplier and data cost. Asking the customer, frozen as W4, failed its test on fresh claims. Against the graded fraud check alone, the second opinion adds +4 claims settled alone, +$84 per 1,000 claims, and +1 wrong decision. These comparisons use the first way of choosing the cut-off, which failed its test on fresh claims, with no payment limit.
Before it is real
Before any of this is real: a test on real claims rather than synthetic ones, timed human tasks, and reviewers tested on planted wrong files before the quick confirmation carries a claim, even for requests.
5Every workflow in full
One card per workflow, each in the same order
Why each one exists, the idea, the part of the machinery it changes, its result and the lesson it left.
The lab publishes its exact thresholds so the evidence can be checked. An insurer running this design would keep its operating thresholds confidential, shared with its regulator, and add random audits, because published cut-offs invite claims tailored to pass them.
decision modellanguage modelcodechanged The heavy outline marks the one part a version changes.
The main chain: each one fixes the last one's flaw
The same fields on every card, in the same order. The strip is the machinery every version shares, each part in its role's colour.
W0
Trust the one big answer
Built · misses the bar
Why it exists
The starting point: the policy as first specified.
The idea
Settle a claim alone only when the decision model confidently answers the broad question, "what do you recommend?"
One broad answer→Fixed rules→Decision rule→Pay · Ask · Person
Reads
The decision model's one recommendation, plus the evidence checks
Learned from data
No: hand-set
Certainty needed
The same for every claim
Models
Decision model only: Jev, asked once and replayed
The planted fraudsent to a personright
In development3.4% settled alone (paid and asked together) · 70.59% right · −$100.0k per 1,000 claims
The exact rule
The engine as built: after the fixed rules, act only if the recommendation is "approve" or "ask for information" at or above the confidence threshold, and the documents, photos, estimate and fraud score all pass.
Lesson. One broad answer is rarely confident, and when it is, it is often wrong. Next: stop asking the broad question.
W1
Combine the small answers
Built · loses money
Why it exists
W0 leaned on one broad answer the decision model is bad at.
The idea
Ignore the recommendation. Code reads the narrow answers (complete? photos match? fraud risk?) and acts only if even the weakest one is strong enough.
Eight answers→Fixed rules→Weakest answer→Pay · Ask · Person
Reads
Eight narrow answers
Learned from data
No: hand-set
Certainty needed
The same for every claim
Models
Decision model only: Jev, replayed
The planted fraudsent to a personright
In development36.6% settled alone (paid and asked together) · 100.00% right · −$76.8k per 1,000 claims
The exact rule
After the fixed rules and the specialist check, the approval confidence is the weakest of the narrow answers; act only if it clears the cut-off chosen for 99.5% right on the training claims.
Lesson. Right almost every time, but far too cautious: one doubtful answer sends a good claim to a person, so it costs more staff time than it saves. Next: weigh the answers instead.
W2
Learn how much to trust each answer
Built · paid a fraud
Why it exists
W1 treated every answer as equally important, and the decision model's probabilities sit 9.6 points from the truth on average on these claims.
The idea
A small statistical model, fitted on past claims, learns a weight for each of the fourteen answers and turns them into one score.
Fourteen answers→Fixed rules→Learned weights→Pay · Ask · Person
Reads
All fourteen answers
Learned from data
Yes: two small logistic models
Certainty needed
The same for every claim
Models
Decision model only: Jev, replayed
The planted fraudpaid automaticallywrong
In development57.8% settled alone (paid and asked together) · 99.83% right · −$39.4k per 1,000 claims
The exact rule
Two logistic models score "paying is right" and "asking for information is right"; the likelier one is proposed, and acted on alone if its score clears the cut-off.
Lesson. Accuracy is not enough: it is right almost every time, but one costly mistake, a planted fraud of $49,101, wipes out every saving. Next: ask for more certainty when more money is at stake.
W3
Stricter as amounts rise
Passed twice
Why it exists
W2 needed the same certainty to pay a small claim or a large one.
The idea
The same learned weights as W2, but above $5,000 the doubt allowed shrinks as the amount grows: at twice that amount, the doubt must be half as large.
Fourteen answers→Fixed rules→Weights × amount→Pay · Ask · Person
Reads
All fourteen answers, and the amount
Learned from data
Yes: the same weights as W2, frozen
Certainty needed
Rises with the amount above $5,000
Models
Decision model only: Jev, replayed; asked afresh in each test on fresh claims
The planted fraudsent to a personright
On fresh claimsThe first version failed. A corrected version passed, then passed a larger, stricter test. The tests in full.
The exact rule
For a payment, the doubt is one minus the score, times the larger of one and the amount divided by $5,000. It pays alone only if one minus the doubt clears the cut-off: 0.813 in the version that passed, chosen on claims its models had not seen. Asking for information pays nothing, so it is never scaled.
Lesson. Holds the accuracy bar on claims it had never seen; it saves money only with its payment limit. Its known weak spot: a few large planted frauds the decision model misreads.
Tried after W3
Each passed in development and earned a test on fresh claims, except W5, which was only ever an experiment.
W3 +
A fifteenth question
Failed on fresh claims
Why it exists
Most of W3's early mistakes were one kind of claim: damage that may predate the incident, which none of the fourteen answers can see.
The idea
Ask the decision model one more question, "could some of the damage predate the incident?", and let W3 weigh the answer.
Fifteen answers→Fixed rules→Decision rule→Pay · Ask · Person
Reads
Fifteen answers
Learned from data
Yes: refitted with the new answer
Certainty needed
As W3
Models
Decision model: Jev, one more question per claim
On fresh claims65.6% settled alone (paid and asked together) · 4 wrong · 99.39% right: below the bar Failed
The exact rule
W3 with the new answer as a fifteenth input: the same regression, risk scaling and cut-off rule.
Lesson. The question caught what it was aimed at, but its answer also made W3 ask other customers for information they did not need, and a wrong request counts against accuracy.
W4
Ask the customer
Failed on fresh claims
Why it exists
Many claims W3 sends to a person are missing one piece of information the customer could give.
The idea
When only the certainty cut-off held a claim back, ask the customer one question, then decide again on the reply.
Fourteen answers→Fixed rules→Decision rule→Asks the customer→Pay · Ask · Person
Reads
W3's answers, then the reply, read again by the decision model
Learned from data
Yes: W3's weights
Certainty needed
As W3; a reply settles at most $20,000
Models
Decision model: Jev. Language model: Claude Haiku 4.5, which writes the question and, in this lab, the customer's reply
On fresh claims59.9% settled alone (paid and asked together) · 1 wrong · net −$43,069 per 1,000: lost money Failed
The exact rule
W3; a claim held back only by the cut-off, with no other block, gets one question. A reply can never lower suspicion, and pays automatically only up to its cap of $20,000.
Lesson. The step that asks the customer made no wrong payment. The loss was a large planted fraud that plain W3 paid by itself, before any question.
W4 +
A payment limit
Failed on fresh claims
Why it exists
W4's one mistake was a large payment.
The idea
Nothing over $20,000 is paid without a person, whichever step proposes it.
Fourteen answers→Rules, with a limit→Decision rule→Asks the customer→Pay · Ask · Person
Reads
As W4
Learned from data
Yes: W3's weights
Certainty needed
As W4, and a person for every payment over $20,000
Models
As W4
On fresh claims52.2% settled alone (paid and asked together) · 1 wrong · net −$3,372 per 1,000: lost money Failed
The exact rule
W4, plus a limit on any automatic payment: above $20,000, a person approves.
Lesson. It stopped the large mistake but sent correct large payments to people, which cost more than it saved. A fault in the run also left some claims unasked. Where to draw that line is a policy choice.
W5
A second opinion
Experiments only
Why it exists
The decision model is unsure of some claims a second reader might settle, and a few large frauds fool it.
The idea
The language model reads the claim beside the decision model's findings, and a claim is settled alone only if both agree.
Fourteen answers→Fixed rules→Decision rule→Second opinion→Pay · Ask · Person
Reads
The claim and the decision model's findings
Learned from data
No
Certainty needed
Agreement, with a certainty floor
Models
Decision model: Jev. Language model: Claude Haiku 4.5, one review per claim
In developmentOn near-misses: 63.0% settled alone (paid and asked together), 99.52% right. On every payment over $20,000 that W3 proposed, it objected to 30 of the 56 W3 would rightly pay, and let 3 of the 20 wrong ones through.
The exact rule
Near-misses: the language model reviews the claims W3 held back only by its cut-off, and a claim is settled on agreement above a floor. Large payments: it reviews every payment over $20,000 that W3 proposes.
Lesson. A negative result, kept on the record: a second vote on the same facts adds little, because it reads the decision model's findings. Never frozen, and never tested on fresh claims.
In the plan, not built
Two ideas from the original plan. Shown so the numbering makes sense, and because W6 is a decision only the business can make.
Plan W3
Ask twice
In the plan, not built
Why it exists
The decision model sometimes gives a different answer to the same question.
The idea
Ask unstable questions again, reworded, and act only when the answers agree.
Asked twice→Fixed rules→Decision rule→Pay · Ask · Person
Reads
The fourteen answers, twice where unstable
Learned from data
No
Certainty needed
Agreement
Models
Decision model: Jev, asked again on some claims
StatusNot built. Steadier decisions, but few more of them, and no help with the large-payment risk.
The exact rule
Planned: re-ask the questions whose answers flip, with reworded instructions, and require agreement before acting alone.
Lesson. Its number went to "stricter as amounts rise".
W6
Policy changes
In the plan, not built
Why it exists
Some claims can only be settled alone if the business changes its rules.
The idea
Automatic denial with a written reason and an appeal path, and a specialist's pre-analysis for complex claims.
Fourteen answers→Fixed rules→Decision rule→Deny, with appeal
Reads
As W3
Learned from data
No
Certainty needed
As W3
Models
Small
StatusNot built: business and regulatory decisions, off by default, and the business's to make.
The exact rule
Planned as switches, each shown with its consequences, never silently on.
Lesson. The plan expected a few points more at most.
6The tests on fresh claims
Tested on fresh claims: it failed, then a corrected version passed
Every other figure on this site is cross-validated on the claims the lab was built with. The real test is claims it has never seen. The lab freezes its deciding step, W3, and runs it once on a thousand fresh claims, under a seed drawn only when the run starts. A result is final for the version tested.
W3, first attempt: failed
In development (cross-validated)
On fresh claims (held-out)
Paid with no person
568 of 1,000
592 of 1,000
Asked for documents, with no person
58 of 1,000
62 of 1,000
Wrong decisions
2 (1 paid, 1 asked)
6 (4 paid, 2 asked)
Accuracy (honest range)
99.68% (98.84% to 99.91%)
99.08% (98.01% to 99.58%)
Wrongly paid
$4,211
$16,194
Net per 1,000 claims
+$5,272
−$6,363
Paid and asked for documents apart: recomputed from the stored answers and checked against the recorded result.
accuracy on handled claims is 99.08% (648 of 654), below 99.50%
net per 1,000 claims is $-6,363, not above $0
Not proven on unseen claims.
Synthetic evidence only; passing this gate is not proof of real-world accuracy and does not promote the configuration.
Priced with the costs recorded when the version was frozen, the lab's earlier placeholders, not the benchmark costs used elsewhere on this site: $55 an hour; a standard review 25 minutes, a specialist 90, a fraud investigation 120, quarantine 45. The accuracy result does not depend on costs.
Margins: accuracy 99.08% against the 99.5% bar, missed by 0.42 points; it needed 3 fewer wrong decisions to pass. Net −$6,363 per thousand claims against $0, missed by $6,363; every claim analysed.
Model: the lab asked for ~typesafe/jev-latest; every answer came from typesafe/jev-1.13-20260917 (4,000 calls).
Which version, run and configuration
Version ec04e73e-87cd-42a2-beda-2db8079a380c; fit 47e0ed13c660d491be0dd4bcdef66fec67e948d52825839068d5f6e53866f825; fit build 2398f73ba3e64619e70ed5c0134b2ab04c26982c.
Held-out run 1f7861a9-6a70-4aeb-a23e-a20a8086160f; slot 86447edd-6344-4eef-8965-d11ead070e7e; result build 2398f73ba3e64619e70ed5c0134b2ab04c26982c.
Synthetic evidence only; passing this gate is not proof of real-world accuracy and does not promote the configuration.
Priced with the costs recorded when the version was frozen, which are the benchmark costs used everywhere else on this site.
Margins: accuracy 100.00% against the 99.5% bar, cleared by 0.50 points; the bar would still have absorbed 2 more wrong decisions. Net +$6,790 per thousand claims against $0, cleared by $6,790; every claim analysed.
Model: the lab asked for ~typesafe/jev-latest; every answer came from typesafe/jev-1.13-20260917 (4,000 calls).
Which version, run and configuration
Version 78a03982-77dc-4e32-9c38-e020bd6157e0; fit cd338fa666922d795df0a3c4f1d144f9c7951226672732d3764866ceeb84ee81; fit build f994cb4157d1b4ffd741ee53b868ddf697a58065.
Held-out run b8ac5efa-f0fc-41e1-b7a3-ae820582a2ff; slot 32621ed9-b6a5-4701-965b-b74ed276b1d4; result build f994cb4157d1b4ffd741ee53b868ddf697a58065.
Decision model: asked for ~typesafe/jev-latest; answered by typesafe/jev-1.13-20260917 (4,000 calls).
Then confirmed on a larger set of fresh claims: passed
A version that passes may take one larger confirmation, on two thousand more fresh claims, declared before its seed is drawn. It is judged more strictly: the bottom of the accuracy range, not the headline figure, must clear the bar. A headline figure can clear the bar by luck on a few hundred claims; the bottom of the range clears it only when the claims are many and the mistakes few. A confirmation never changes the test result it follows.
On 2,000 more fresh claims
Paid with no person
1,095 of 2,000
Asked for documents, with no person
95 of 2,000
Wrong decisions
1 (1 paid, 0 asked)
Accuracy (honest range)
99.92% (99.53% to 99.99%)
Bottom of the range, judged against the bar
99.53% against 99.5%
Wrongly paid
$954
Net per 1,000 claims
+$12,613
Paid and asked for documents apart: recomputed from the stored answers and checked against the recorded result.
Held-out confirmation on synthetic claims: passed. 2000 of 2000 analysed; 0 unresolved. Complete.
Stored rule: verdict/2; the lower end of the 95% accuracy range >= 0.995, net per 1,000 > 0.
Synthetic evidence only; a confirmation neither changes the gate result nor promotes the configuration.
Priced with the costs recorded when the version was frozen, which are the benchmark costs used everywhere else on this site.
Margins: the bottom of the accuracy range 99.53% against the 99.5% bar, cleared by 0.03 points; one more wrong decision would have failed it. Net +$12,613 per thousand claims against $0, cleared by $12,613; every claim analysed.
Model: the lab asked for ~typesafe/jev-latest; every answer came from typesafe/jev-1.13-20260917 (8,000 calls).
Which confirmation run
Confirmation run f559ede7-133b-4688-b202-ea1e13177b9f; slot f4d5cc1a-e7f3-4096-8f4d-4d872a4a25cb; result build 02f965d7726b1d4fb227df250913366ed1ce14e4.
Decision model: asked for ~typesafe/jev-latest; answered by typesafe/jev-1.13-20260917 (8,000 calls).
A test declared in advance, with the payment limit: passed
Declared before it ran, with a public timestamp, for the version as it runs now, payment limit included, on fresh claims drawn after the limit existed; the declaration itself is published with the result. It is judged on the bottom of the accuracy range, and it changes neither the version's result nor its standing.
To check it: the declaration named drand round 32600518, and its timestamp is in Bitcoin block 968849. The declaration, its proof and the round's beacon are in the published record. What this proves: each test's version and settings were declared before its seed could be known. The declaration is in a Bitcoin block at least two hours before the public drand round it names, and the seed is computed from that round. The log is hash-chained and lists every declared test with its status, so a test declared and then dropped still shows. What it can't prove: declarations newer than the latest published one could still be held back, and the block header is checked offline, so confirm its height in any Bitcoin explorer.
On 3,000 fresh claims
Paid with no person
1,486 of 3,000
Asked for documents, with no person
127 of 3,000
Wrong decisions
1 (1 paid, 0 asked)
Accuracy (honest range)
99.94% (99.65% to 99.99%)
Bottom of the range, judged against the bar
99.65% against 99.5%
Wrongly paid
$3,293
Net per 1,000 claims
+$4,059
Paid and asked for documents apart: recomputed from the stored answers and checked against the recorded result.
Held-out config-check on synthetic claims: passed. 3000 of 3000 analysed; 0 unresolved. Complete.
Stored rule: verdict/2; the lower end of the 95% accuracy range >= 0.995, net per 1,000 > 0.
Run under the configuration that decided when it ran, with no automatic payment above $20,000.
Synthetic evidence only; a config-check changes neither the version's status nor its promotion.
Priced with the costs recorded when the version was frozen, which are the benchmark costs used everywhere else on this site.
Margins: the bottom of the accuracy range 99.65% against the 99.5% bar, cleared by 0.15 points; the bar would still have absorbed 1 more wrong decision. Net +$4,059 per thousand claims against $0, cleared by $4,059; every claim analysed.
Model: the lab asked for ~typesafe/jev-latest; every answer came from typesafe/jev-1.13-20260917 (12,000 calls).
Which run
Config-check run fa1c09db-78ac-4527-a9e4-d50c8e2bbc85; slot 78456d61-116c-4a64-a67f-dc039ebe4d8b; result build 34060d27e272276ac3f989cd447521a36e4a84d2.
Decision model: asked for ~typesafe/jev-latest; answered by typesafe/jev-1.13-20260917 (12,000 calls).
What changed between the attempts: only how the certainty cut-off is chosen. The first version picked the lowest cut-off that was right often enough on the claims it had just learned from, and a model is always more sure of claims it has seen. The second version scores those same claims with models fitted without them, and picks its cut-off from those scores. The failure prompted the change, but the new cut-off was chosen on the lab's own claims, never on the held-out ones, and the second version was tested on a new set of fresh claims.
The corrected version made no wrong decisions and a profit, though it settles fewer claims alone. These are synthetic claims: a pass starts the evidence, it does not end it.
Each result is final for its frozen version: the first failure stays on the record. A second attempt needed a new frozen version, a newly declared test slot and a new spending approval. The declarations for the tests so far are kept in the lab's private records. From here on, each test's declaration, including the frozen version's hash, is timestamped publicly before the test runs. The first test declared this way has run: its declaration was timestamped in a Bitcoin block before the public drand beacon round it named, and its fresh claims were drawn from that round.
It then passed a larger, stricter test: the bottom of its range, 99.53%, cleared the 99.5% bar.
The development figures for every version still use the first way of choosing the cut-off, except the recommended design, which has been replayed the corrected way, with the payment limit. The version that passed is stricter, so it settles fewer claims with no person than the development W3 shows.
The designs built after W3
Three designs were frozen after W3 on the same development claims, each with its own test on a thousand fresh claims. All three failed. Their records are kept apart from W3's, and none changes W3's result.
W3 + a fifteenth question Failed
Asks the decision model whether the damage could predate the incident: the one kind of claim that caused most of W3's early mistakes. The question caught what it was aimed at. But its answer also made W3 ask other customers for information they did not need, and a wrong request counts against accuracy.
accuracy on handled claims is 99.39% (652 of 656), below 99.50%
Not proven on unseen claims.
Synthetic evidence only; passing this gate is not proof of real-world accuracy and does not promote the configuration.
Priced with the costs recorded when the version was frozen, which are the benchmark costs used everywhere else on this site.
Margins: accuracy 99.39% against the 99.5% bar, missed by 0.11 points; it needed 1 fewer wrong decision to pass. Net +$13,196 per thousand claims against $0, cleared by $13,196; every claim analysed.
Model: the lab asked for ~typesafe/jev-latest; every answer came from typesafe/jev-1.13-20260917 (5,000 calls).
Which version, run and configuration
Version 37e208ce-d730-49ce-af47-9e8126076dfc; fit ebcb4ac8ae858cc9f1e96862939d91de8293f406981aecf051a25f13acc9a605; fit build 33b0bcbecfda2584d875732cf6bc4099849ddd36.
Held-out run dc4f870c-f86f-4e3a-b3ad-40bd3dd79044; slot afc69a57-644f-4766-8ac7-5ab703c5d0b1; result build 33b0bcbecfda2584d875732cf6bc4099849ddd36.
Decision model: asked for ~typesafe/jev-latest; answered by typesafe/jev-1.13-20260917 (5,000 calls).
W4: ask the customer Failed
When one piece is missing, ask the customer for it, then decide again. A reply can settle a claim only up to a cap. The step that asks the customer made no wrong payment, and caught every attempt to manipulate it, against five fixed test phrasings; subtler attacks are not yet tested. The claim that failed the test was one plain W3 paid by itself.
the evidence is incomplete: not every claim was analysed, and an incomplete run cannot pass
Not proven on unseen claims.
Synthetic evidence only; passing this gate is not proof of real-world accuracy and does not promote the configuration.
Priced with the costs recorded when the version was frozen, which are the benchmark costs used everywhere else on this site.
Margins: accuracy 99.83% against the 99.5% bar, cleared by 0.33 points; the bar would still have absorbed 1 more wrong decision. Net −$43,069 per thousand claims against $0, missed by $43,069; only 999 of 1,000 claims analysed, short of the complete run the rule requires.
Model: the lab asked for ~typesafe/jev-latest; every answer came from typesafe/jev-1.13-20260917 (4,624 calls).
Which version, run and configuration
Version 82d30d60-d225-4047-9779-58816c46de28; fit 82640130a404c5e392a208abb355390b46ca147ff5c866a3604551f71b44127c; fit build 78c42acac3cb24c5d544f5444e0504638d65f261.
Held-out run 99e4ee1a-9911-469b-95b9-270752bbd986; slot 2ecf46bf-333c-4e11-8b4b-df1062936df8; result build 78c42acac3cb24c5d544f5444e0504638d65f261.
Decision model: asked for ~typesafe/jev-latest; answered by typesafe/jev-1.13-20260917 (4,624 calls).
W4 + a payment limit Failed
Nothing over the limit is paid without a person, whichever step proposes it. The limit stopped the large mistake, but sent correct large payments to a person, which cost more staff time than it saved. A fault in the run also left some claims unasked, so its evidence is incomplete.
the evidence is incomplete: not every claim was analysed, and an incomplete run cannot pass
Not proven on unseen claims.
Synthetic evidence only; passing this gate is not proof of real-world accuracy and does not promote the configuration.
Priced with the costs recorded when the version was frozen, which are the benchmark costs used everywhere else on this site.
Margins: accuracy 99.81% against the 99.5% bar, cleared by 0.31 points; the bar would still have absorbed 1 more wrong decision. Net −$3,372 per thousand claims against $0, missed by $3,372; only 905 of 1,000 claims analysed, short of the complete run the rule requires.
Model: the lab asked for ~typesafe/jev-latest; every answer came from typesafe/jev-1.13-20260917 (4,284 calls).
Which version, run and configuration
Version 9beb673c-403a-4afc-818d-2f0546085370; fit afb2340e5fee7d98410a679616c0498d47b4c7d5995b24a19c8ff8712725d35a; fit build 5a0a26c0a5360e2a1e1615ca062c81025b1242da.
Held-out run 59468f88-75c3-4ff3-8905-768ac1650f81; slot 392f1438-733f-4419-ae53-b3f3f89f3bbc; result build 5a0a26c0a5360e2a1e1615ca062c81025b1242da.
Decision model: asked for ~typesafe/jev-latest; answered by typesafe/jev-1.13-20260917 (4,284 calls).
Every set of fresh claims, pooled
The version that passed was also replayed, frozen and never refitted, on the fresh claims drawn for the other tests: the first W3 attempt's and the three later designs'. Its own test and confirmation replay exactly to what they recorded. The version in shadow mode adds one rule: nothing over $20,000 is paid without a person.
With the payment limit, as it runs in shadow mode
Fresh claims from
Claims
Paid alone
Asked for documents
Wrong
Net per 1,000
Range
W3, first attempt
1,000
49.9%
4.4%
1
+$4,410
−$5,250 to +$13,370
Its own test (W3, second attempt)
1,000
50.7%
4.4%
0
+$3,850
−$5,740 to +$13,370
Its confirmation
2,000
50.7%
4.8%
1
+$9,813
+$3,465 to +$15,680
W3 + a fifteenth question
1,000
48.2%
3.6%
0
+$8,820
$0 to +$16,660
W4: ask the customer
1,000
47.4%
4.6%
0
+$1,680
−$8,050 to +$10,500
W4 + a payment limit
1,000
47.3%
4.0%
1
−$1,692
−$12,612 to +$8,120
Its test declared in advance, with the payment limit
3,000
49.5%
4.2%
1
+$4,059
−$1,729 to +$9,519
All sets, pooled
10,000
49.4%
4.3%
4
+$4,887
+$1,861 to +$7,951
Without the limit, as it was tested
Fresh claims from
Claims
Paid alone
Asked for documents
Wrong
Net per 1,000
Range
W3, first attempt
1,000
55.0%
4.4%
1
+$7,980
−$1,470 to +$17,150
Its own test (W3, second attempt)
1,000
54.9%
4.4%
0
+$6,790
−$2,730 to +$15,750
Its confirmation
2,000
54.8%
4.8%
1
+$12,613
+$6,361 to +$18,563
W3 + a fifteenth question
1,000
54.0%
3.6%
0
+$12,880
+$4,340 to +$21,000
W4: ask the customer
1,000
53.2%
4.6%
1
−$43,699
−$147,827 to +$12,390
W4 + a payment limit
1,000
53.1%
4.0%
1
+$2,368
−$8,214 to +$12,518
Its test declared in advance, with the payment limit
3,000
54.9%
4.2%
1
+$7,792
+$2,449 to +$13,300
All sets, pooled
10,000
54.4%
4.3%
5
+$3,492
−$8,139 to +$10,728
Net ranges are the middle ninety-five percent of 2,000 seeded resamples of the claims, at the benchmark staff costs. They reflect how often large planted frauds occur in these sets of fresh claims, not how often they would occur elsewhere.
With no model, on rules alone, the same W3 method settles 4.7% of the development claims with no person and nets +$3,290 per thousand claims; with the decision model, 62.6% and +$8,809.
Accuracy against the share settled, if the line moved
How often W3 is right as its line is moved, on 10,000 pooled fresh claims and, dashed, on the development claims. The version keeps the line it was frozen with; nothing here chooses another.
Accuracy against the share of claims settled with no personSolid: the fresh claims, with the honest range shaded. Dashed: the development claims. The dot is W3's frozen line. The share axis starts at 30%: below it so few claims are settled that their accuracy says little.
W3's frozen line, a certainty cut-off of 0.813, settles 53.7% of these fresh claims with no person (paid and asked together), 99.93% right (range 99.81% to 99.97%).
The most W3 could settle with no person while keeping each accuracy target, and at which cut-off: by the measured accuracy, and by the bottom of its range.
Accuracy target
Fresh claims, by the measured accuracy
Fresh claims, by the bottom of the range
Development claims, by the measured accuracy
99.0%
58.5% (cut-off 0.575)
57.9% (cut-off 0.626)
58.7% (cut-off 0.615)
99.5%
57.0% (cut-off 0.689)
56.3% (cut-off 0.720)
57.8% (cut-off 0.649)
99.9%
54.8% (cut-off 0.775)
52.9% (cut-off 0.834)
53.5% (cut-off 0.780)
Descriptive only. This curve is drawn on held-out claims after the fact; a cut-off chosen from it would be fitted to them, so no cut-off is chosen here, and the version keeps the one it was frozen with. The payment limit was itself set after one of these sets (see the retrospective note), so this curve inherits that.
Only W3's line moves: every hard rule, and the payment limit, stays exactly as frozen at every point.
The development curve is development evidence: the same design cross-validated on the evaluation claims, the claims the designs were built on. Where the two curves part is how far development flattered the design.
7Costs and the people it takes
The people it takes
Staff time, turned into the number a business plans around: people. Each figure is one full-time person doing claim work.
ModelChange the assumptions
How many people each version needs
The evidence counts the staff work behind every version, task by task: 122.0 min of staff time per claim by hand, 91.4 min in the recommended version, audits included. Change any assumption below; the work stays what the evidence measured, only its price and the headcount move.
By hand
Recommended version
Same team, more claims
Where the people go
Standard review
Quick confirm
Light fraud check
Fraud investigation
Specialist review
Quarantine
Audits after the decision
By hand Recommended
Where the people go
By hand
Recommended
Standard review
Quick confirm
Light fraud check
Fraud investigation
Specialist review
Quarantine
Audits after the decision
Claims per year
People by hand
People, recommended
Saved a year
10,000
100,000
1,000,000
Full-time equivalents doing claim work only: supervisors, managers and queue slack would make both columns larger. The recommended version's fraud investigations do not shrink: they are the claims that genuinely need one, and after automation they are most of the remaining work.
Automation moves people to the claims that need them
The recommended design does not remove people from claims. It takes routine claims off their desks: the clean ones are settled with no person, and a request for more information the model is nearly sure of gets a quick confirmation instead of a full review. Every approval a person must see gets a full review.
The same team can take on more claims before it has to grow.
Where the cost assumptions come from
Assumption
Used here
Benchmark
Source
Standard review
60 minutes
30 to 90 minutes; simple home claims take 1 to 3 hours On its test on 2,000 fresh claims, W3 still saves money while this takes at least 43.3 min; nothing published yet confirms or rules that out.
Industry practice; no published touch-time measurement (weak)
Specialist review
6 hours
4 to 10 hours or more for complex files
Industry practice (weak)
Fraud investigation
9 hours
174 accepted cases per investigator per year: about 9 to 10 hours each
Decision within 21 days of proof of loss (regulation); final payment on home claims about 41 days. Five days is likely optimistic for people, and no source measures the time to a decision on routine claims
A pass on fresh synthetic claims shows the result reproduces on claims from the same generator, not that it carries over to real ones. The recommended design has been replayed with the corrected way of choosing the cut-off; the other steps built around W3 have not. The evidence adds:
A frozen W3 has held-out evidence; every other figure here is development evidence on synthetic claims.
Staff minutes per task and the $70 cost of a productive hour are benchmark mid-points, and the short lanes' and audits' minutes are the lab's assumptions: none are measurements.
A person is assumed to decide every claim they see correctly; people's own errors are not measured.
An automated request for information is a next action only: the customer's reply and the claim's final outcome are not recorded.
The second opinion rests on real calls for 110 near-miss claims only.
Modeled net is staff capacity at assumed costs, not realized cash savings; this is a demonstration, not a customer deployment.
The 99.5% accuracy bar is the lab's starting assumption, set when the lab began and taken from no outside source: it is not an industry standard. Published figures on people's accuracy measure other things, and people's own accuracy on these claims is not measured.
9What the research says
What the research says
Each way the lab tests a design guards against a risk that published research names, and the first W3's failure is what that research predicts. The same research also cuts against two of the steps built around W3, and names the limit that matters most.
Where the method follows it
Choosing the cut-off on claims the models had not seen
The Wilson interval stays reliable near perfect accuracy, where the textbook interval does not. The stricter test judges the bottom of that range. Brown, Cai & DasGupta, 2001, Statistical Science
Money as well as accuracy
Counting mistakes alone misleads when mistakes differ in cost: W2 is right almost every time and still loses money. Elkan, 2001, IJCAI
A review of more than a hundred experimental studies found that, on decision tasks, a person working with an AI did worse on average than the better of the two alone, most of all when the automated system alone was the stronger. A person approving a prepared file has that shape. Its saving stays a hypothesis until reviewers are tested with planted wrong files. Vaccaro, Almaatouq & Malone, 2024, Nature Human Behaviour
The second opinion
Combining two deciders pays when they are diverse, when their mistakes differ. When two language models both err, they often give the same wrong answer, and more accurate models more so, and this one also reads the decision model's findings. A second vote on the same facts is weak evidence: the certainty floor does the work, and the lab keeps it as a negative result. A language model is worth its cost where it reads or checks something the decision model cannot. Steyvers, Tejeda, Kerrigan & Smyth, 2022, PNAS; Kim, Garg, Peng & Garg, 2025, ICML
1,000 synthetic claims, answered by ~typesafe/jev-latest through jev-openrouter. 5-fold cross-validation: every claim is scored by a workflow whose cut-offs and weights were fitted on the other four folds.
The claims are generated by code from a fixed seed, across auto, home, life and commercial insurance; the generator sets each claim's right answer and plants the frauds, manipulation attempts and missing papers.
Disclosure: the lab has no relationship with TypeSafe, the company that makes Jev. Jev was used through OpenRouter, and the lab's owner paid for every model call, Jev's and the language models', personally. The lab's owner designed, ran and judged every test, working with AI coding assistants; nothing here has yet been replicated independently.
Evidence: run 6827826e-b962-4bd8-bdb3-b163712fd660, exported from build 7c88683e4bce, dataset 96f9d5e30b51. Synthetic claims and stated assumptions: a demonstration, not a customer deployment.
Costs are benchmark mid-points, stated once
Staff minutes per task and the cost of a productive hour are mid-points of published industry benchmarks, listed below with their sources. Four have no benchmark and are the lab's own assumptions. None are any insurer's records, so read the dollar figures for direction and relative size, not as a savings forecast.
How each number is known, strongest first
Held-out testA frozen version run once on fresh claims it never saw, under a seed drawn when the run started. The strongest evidence here, and the only kind that can pass or fail the design.W3, 2 frozen versions
Tested with real callsA real second model was called on the claims it would see.+2nd
Measured in the engineThe rules as first built, run once over every claim. Measured differently from the later steps, so part of the jump to W1 is the method.W0
Cross-validatedEach claim is graded by settings tuned on the other claims, never on itself.W1, W2, W3
Replayed, assumed minutesThe stored answers routed through the new process for real; the minutes each new task takes are assumed.+QC, +Fraud, +Audit
Simulated repliesCustomer replies written by a language model, including deliberately misleading ones.+Ask
Assumed starting pointThe by-hand process, priced with the lab's assumed minutes. Nothing is measured.By hand
The evidence behind every figure on this site is published as it was exported, byte for byte: the manifest names each file and its hash.