On a test declared in advance, code and a narrow decision model paid 49.5% of fresh synthetic claims with no person and asked for documents on 4.2%, right on 99.94% of the claims it decided.

The test: the version in shadow mode, with its payment limit, on 3,000 fresh claims drawn after the limit, declared before it ran. It is one test; every other test and every set pooled are below. A decision model, Jev, answers narrow questions, code makes the decision, people take the rest. The claims are synthetic, and nothing here is validated in real use.

Figures last updated , from the published evidence. What changed

The short version

49.5%paid with no person (1,486 of 3,000 fresh claims; range 47.7% to 51.3%), by the version in shadow mode, which pays nothing automatically over $20,000

4.2%asked for documents, with no person (127 of 3,000; range 3.6% to 5.0%): a request is a next step, not a resolved claim, and is never added to the paid share

1wrong decision (1 paid, 0 asked for documents) in the 1,613 decided with no person: 99.94% right (range 99.65% to 99.99%), judged on the bottom of the range against a 99.5% bar

+$4,059net per thousand claims on this test, drawn after the limit (range −$1,729 to +$9,519). Its range crosses zero: one set cannot show the money is positive.

How the test was declared, and how to check it

Beside it: every set of fresh claims, pooled

Every set the version can be scored on, the test above included. This pooled money figure looks back.

49.4%paid with no person (4,936 of 10,000 fresh claims), pooled over seven sets, by the version in shadow mode, which pays nothing automatically over $20,000

4.3%asked for documents, with no person (432 of 10,000): a request is a next step, not a resolved claim, and is never added to the paid share (each share is rounded on its own)

4wrong decisions (3 paid, 1 asked for documents) in the 5,368 decided with no person: 99.93% right (range 99.81% to 99.97%), against a 99.5% bar

+$4,887net per thousand claims, pooled: a retrospective analysis (range +$1,861 to +$7,951)

Each wrong decision, in W3's own steps

The payment limit was set after W4's test made a wrong payment of $49,999, so the pooled money figure looks back: five of the seven sets were drawn before the limit existed. On sets drawn after it was set, the version in shadow mode nets −$1,692 per thousand claims (−$12,612 to +$8,120) on the first, and +$4,059 (−$1,729 to +$9,519) on a test declared in advance. Without the limit the pooled range crosses zero.

Every test on fresh claims

Every test on fresh claims: 3 passed and 4 failed. The figures are in each row.

The design that passed: W3, stricter as amounts rise

W3, first attempt1,000 fresh claims

65.4% settled alone: 59.2% paid, 6.2% asked · 6 wrong (4 paid, 2 asked)

Failed99.08% right: below the bar

Margins: accuracy 99.08% against the 99.5% bar, missed by 0.42 points; it needed 3 fewer wrong decisions to pass. Net −$6,363 per thousand claims against $0, missed by $6,363; every claim analysed.

W3, second attempt1,000 fresh claims

59.3% settled alone: 54.9% paid, 4.4% asked · 0 wrong (0 paid, 0 asked)

Passed100.00% right, and saves money

Margins: accuracy 100.00% against the 99.5% bar, cleared by 0.50 points; the bar would still have absorbed 2 more wrong decisions. Net +$6,790 per thousand claims against $0, cleared by $6,790; every claim analysed.

W3, second attempt: confirmation2,000 more, judged more strictly

59.5% settled alone: 54.8% paid, 4.8% asked · 1 wrong (1 paid, 0 asked)

Passedbottom of its range 99.53%

Margins: the bottom of the accuracy range 99.53% against the 99.5% bar, cleared by 0.03 points; one more wrong decision would have failed it. Net +$12,613 per thousand claims against $0, cleared by $12,613; every claim analysed.

W3, second attempt: declared in advance, with the payment limit3,000 fresh claims, drawn after the limit

53.8% settled alone: 49.5% paid, 4.2% asked · 1 wrong (1 paid, 0 asked)

Passedbottom of its range 99.65%

Margins: the bottom of the accuracy range 99.65% against the 99.5% bar, cleared by 0.15 points; the bar would still have absorbed 1 more wrong decision. Net +$4,059 per thousand claims against $0, cleared by $4,059; every claim analysed.

Attempts to go further

W3 + a fifteenth question1,000 fresh claims

65.6% settled alone (paid and asked together) · 4 wrong

Failed99.39% right: below the bar

Margins: accuracy 99.39% against the 99.5% bar, missed by 0.11 points; it needed 1 fewer wrong decision to pass. Net +$13,196 per thousand claims against $0, cleared by $13,196; every claim analysed.

W4: ask the customer1,000 fresh claims

59.9% settled alone (paid and asked together) · 1 wrong

Failednet −$43,069 per 1,000: lost money

Margins: accuracy 99.83% against the 99.5% bar, cleared by 0.33 points; the bar would still have absorbed 1 more wrong decision. Net −$43,069 per thousand claims against $0, missed by $43,069; only 999 of 1,000 claims analysed, short of the complete run the rule requires.

W4 + a payment limit1,000 fresh claims

52.2% settled alone (paid and asked together) · 1 wrong

Failednet −$3,372 per 1,000: lost money

Margins: accuracy 99.81% against the 99.5% bar, cleared by 0.31 points; the bar would still have absorbed 1 more wrong decision. Net −$3,372 per thousand claims against $0, missed by $3,372; only 905 of 1,000 claims analysed, short of the complete run the rule requires.

Share of all claims settled with no person, paid and asked for documents together, on claims each version had never seen. A test passes only if enough of those decisions are right and the version saves money after paying for its mistakes. Where a row shows paid and asked apart, they were recomputed from the stored answers and checked against the recorded result. The accuracy bar is 99.5% right. Each track runs to the ceiling, 78.4%: the most that could ever be settled alone.

In the lab's baseline, a person reviews every claim, even the routine ones.

Today an adjuster reads every claim, checks the policy by hand and decides. That is about 122.0 min of staff time per claim and about 5.0 days of waiting, most of it in a queue, though most claims are routine.

Today. Of 6 steps, people do 5, code 1, the decision model none.
  1. Code: Claim arrives
  2. People: Waits in a queue
  3. People: Adjuster reads everything
  4. People: Adjuster checks the policy by hand
  5. People: Adjuster decides
  6. People: Payment is set up

Where the 1,000 claims end up:

All 1,000 decided by a person, at step 6

The question

Can a narrow, cheap decision model settle routine claims with no person, if it only answers questions and code makes every decision?

What counts as success

  • At least 99.5% of the decisions it makes alone are right
  • Cheaper than doing it by hand, after paying for every mistake
  • Holds on claims it has never seen
  • As many claims settled alone as possible: at most 78.4% can be, because the rest must reach a person by rule

What stays with people

  • Every claim over the large-claim limit
  • Every denial: none is automatic
  • Signs of fraud, and any attempt to manipulate the model

The decision model answers questions. Code makes every decision.

The decision model is never asked what to do with a claim. It answers fourteen narrow questions, each with a probability: is the claim complete, do the photos match, how strong are the fraud signs. Code turns the answers into a decision: pay, ask the customer for what is missing, or send the claim to a person.

The decision model here is Jev, TypeSafe's first System One model, trained with RLCD: reinforcement learning for calibrated decisions. Jev is TypeSafe's model; the lab calls it through OpenRouter and changes nothing about it. A language model earns its cost only where it does what the decision model cannot: write, read a claim's documents, or check a reply against them. Changing the models, and the process around them, is how the lab works to raise accuracy and speed.

With no model, on rules alone, the same W3 method settles 4.7% of the development claims with no person and nets +$3,290 per thousand claims; with the decision model, 62.6% and +$8,809.

The design: Jev answers, code decides, people review. Of 11 steps, code does 8, the decision model 1, people 2.
  1. Code: Claim arrives
  2. Decision model: Jev answers narrow questions
  3. Code, a fixed rule: Over $50,000?
  4. Code, a fixed rule: Manipulation attempt?
  5. Code, a fixed rule: Fraud signs: how strong?
  6. Code, a fixed rule: Policy in force?
  7. Code: Certain enough for the amount?
  8. Code: Over $20,000?
  9. People: Quick confirm, requests only
  10. Code: Pay, or ask for what is missing
  11. People: Audit, afterwards

Where claims go:

a person, with the analysis attached, leaving at step 3
quarantine, leaving at step 4
a light check or an investigation, leaving at step 5
a person, who confirms any denial, leaving at step 6
a standard review by a person, leaving at step 7
a person approves the payment, leaving at step 8
paid, or asked for documents, with no person, at step 10

Steps marked as fixed rules run before anything else. No answer from the decision model can change them. The lab publishes its exact thresholds so the evidence can be checked. An insurer running this design would keep its operating thresholds confidential, shared with its regulator, and add random audits, because published cut-offs invite claims tailored to pass them.

Where the claims go under the recommended design Each bar is every claim counted, by the route it took. The figures are in the table below.

Development claims, each scored by settings fitted on the others 1,000 claims, with the payment limit

Each fold's cut-off chosen on claims its models had not seen: the corrected way, as in the version that passed.

The payment limit sends 56 of these to a person.

Every set of fresh claims, pooled 10,000 claims, with the payment limit

The payment limit sends 507 of these to a person.

RouteDevelopment claimsFresh claims, pooled
Paid with no person487 (48.7%)4,936 (49.4%)
Asked for documents, with no person50 (5.0%)432 (4.3%)
A quick confirmation by a person, for requests only36 (3.6%)296 (3.0%)
A light fraud check by a person78 (7.8%)735 (7.3%)
A person: not certain enough for the amount107 (10.7%)1,238 (12.4%)
A person, by a fixed rule or the payment limit242 (24.2%)2,363 (23.6%)

The recommended design on the fresh claims: the frozen W3 with the payment limit decides which claims are paid, asked for more or sent to a person, exactly as the fresh figures are scored. The quick-confirmation floor and the fraud grade were fitted once on the 1,000 evaluation claims and applied unchanged to the fresh claims, which they were never fitted on. Both were fitted with the payment limit, as the design runs.

Decision model: Jev (TypeSafe, System One) beside the language model
MeasureDecision model: Jev (TypeSafe, System One)Language model: Claude Haiku 4.5
Version used~typesafe/jev-latest (answered by typesafe/jev-1.13-20260917)anthropic/claude-haiku-4.5
Trained forCalibrated decisions (RLCD)Answers people prefer, or that can be checked (RLHF, RLVR)
What it returnsTyped values, each with a probability; nothing to parseText, which code must parse and check
Speed hereA median of 276 ms a call; one call in twenty took longer than 515 msNot measured in this lab
Cost here$0.12 per thousand claims, all fourteen questions$0.81 per thousand claims for a second opinion; $1.92 for a question and its reply
Its job hereAnswers the fourteen questions on every claimWrites the question to the customer and, in this lab, the customer's reply. Proposed, not built: reading a claim's documents, and checking a reply against them, with code checking what it reads
What the lab foundIts yes-or-no probabilities sit 9.6 points from how often the answer was yes, on average. Mostly too cautious: when it said 83%, the answer was yes 94% of the timeAs a second vote on the same facts it added little: asked on 110 claims, it let 4 more be settled alone. Kept as a negative result. A made-up reassuring reply once got a planted fraud paid

How the lab runs, and every step in full

Ten ideas. Four designs faced fresh claims, in seven tests; one passed.

Every workflow keeps the same machinery: the decision model answers questions, fixed rules run first, and a decision rule turns the answers into pay, ask or a person. Each one changes one part of it. The main chain fixed one flaw at a time until W3 held up on fresh claims.

THE MAIN CHAIN: EACH ONE FIXES THE LAST ONE'S FLAWTRIED AFTER W3IN THE PLAN, NOT BUILT use thesmall answerslearn theweightsscale bythe money W0 Trust the one biganswer Built · misses the bar W1 Combine the smallanswers Built · loses money W2 Learn how much totrust each answer Built · paid a fraud W3 Stricter as amountsrise Passed twice W5 A second opinion + LANGUAGE MODELExperiments only W3 + A fifteenth question Failed on fresh claims W4 Ask the customer + LANGUAGE MODELFailed on fresh claims W4 + A payment limit + LANGUAGE MODELFailed on fresh claims Plan W3 Ask twice In the plan, not built W6 Policy changes In the plan, not built

The main chain: each one fixes the last one's flaw

  1. W0Trust the one big answerBuilt · misses the bar
  2. W1Combine the small answersBuilt · loses money
  3. W2Learn how much to trust each answerBuilt · paid a fraud
  4. W3Stricter as amounts risePassed twice

Tried after W3

  1. W3 +A fifteenth questionFailed on fresh claims
  2. W4Ask the customer + Language modelFailed on fresh claims
  3. W4 +A payment limit + Language modelFailed on fresh claims
  4. W5A second opinion + Language modelExperiments only

In the plan, not built

  1. Plan W3Ask twiceIn the plan, not built
  2. W6Policy changesIn the plan, not built
  • Passed on fresh claims
  • Failed on fresh claims
  • Built, measured in development
  • Experiment, never frozen
  • Planned, not built

Every workflow uses the decision model, Jev. Those marked + Language model also use the language model, Claude Haiku 4.5.

Why the numbers don't follow the plan. The plan named W0 to W6. Its W3, "ask twice", was never built, and the number went to "stricter as amounts rise", an idea that came from W2 paying a planted fraud. The plan's W4 was built as a frozen workflow; its W5 was tried only as experiments.

Around W3, not workflows: The quick confirmation, the graded fraud check and the audit keep W3's decision and change the human work around it. They are on the process diagram, not in this tree.

The main chain, one line each; open any for its card

W0Trust the one big answerSettle a claim alone only when the decision model confidently answers the broad question, "what do you recommend?"3.4% alone · misses the bar
Why it exists
The starting point: the policy as first specified.
Reads
The decision model's one recommendation, plus the evidence checks
Learned from data
No: hand-set
Certainty needed
The same for every claim
Models
Decision model only: Jev, asked once and replayed

The planted fraud sent to a person right

In development 3.4% settled alone (paid and asked together) · 70.59% right · −$100.0k per 1,000 claims

The exact rule

The engine as built: after the fixed rules, act only if the recommendation is "approve" or "ask for information" at or above the confidence threshold, and the documents, photos, estimate and fraud score all pass.

Lesson. One broad answer is rarely confident, and when it is, it is often wrong. Next: stop asking the broad question.

W1Combine the small answersIgnore the recommendation. Code reads the narrow answers (complete? photos match? fraud risk?) and acts only if even the weakest one is strong enough.36.6% alone · loses money
Why it exists
W0 leaned on one broad answer the decision model is bad at.
Reads
Eight narrow answers
Learned from data
No: hand-set
Certainty needed
The same for every claim
Models
Decision model only: Jev, replayed

The planted fraud sent to a person right

In development 36.6% settled alone (paid and asked together) · 100.00% right · −$76.8k per 1,000 claims

The exact rule

After the fixed rules and the specialist check, the approval confidence is the weakest of the narrow answers; act only if it clears the cut-off chosen for 99.5% right on the training claims.

Lesson. Right almost every time, but far too cautious: one doubtful answer sends a good claim to a person, so it costs more staff time than it saves. Next: weigh the answers instead.

W2Learn how much to trust each answerA small statistical model, fitted on past claims, learns a weight for each of the fourteen answers and turns them into one score.57.8% alone · paid a fraud
Why it exists
W1 treated every answer as equally important, and the decision model's probabilities sit 9.6 points from the truth on average on these claims.
Reads
All fourteen answers
Learned from data
Yes: two small logistic models
Certainty needed
The same for every claim
Models
Decision model only: Jev, replayed

The planted fraud paid automatically wrong

In development 57.8% settled alone (paid and asked together) · 99.83% right · −$39.4k per 1,000 claims

The exact rule

Two logistic models score "paying is right" and "asking for information is right"; the likelier one is proposed, and acted on alone if its score clears the cut-off.

Lesson. Accuracy is not enough: it is right almost every time, but one costly mistake, a planted fraud of $49,101, wipes out every saving. Next: ask for more certainty when more money is at stake.

W3Stricter as amounts riseThe same learned weights as W2, but above $5,000 the doubt allowed shrinks as the amount grows: at twice that amount, the doubt must be half as large.Passed twice
Why it exists
W2 needed the same certainty to pay a small claim or a large one.
Reads
All fourteen answers, and the amount
Learned from data
Yes: the same weights as W2, frozen
Certainty needed
Rises with the amount above $5,000
Models
Decision model only: Jev, replayed; asked afresh in each test on fresh claims

The planted fraud sent to a person right

On fresh claims The first version failed. A corrected version passed, then passed a larger, stricter test. The tests in full.

The exact rule

For a payment, the doubt is one minus the score, times the larger of one and the amount divided by $5,000. It pays alone only if one minus the doubt clears the cut-off: 0.813 in the version that passed, chosen on claims its models had not seen. Asking for information pays nothing, so it is never scaled.

Lesson. Holds the accuracy bar on claims it had never seen; it saves money only with its payment limit. Its known weak spot: a few large planted frauds the decision model misreads.

Follow one claim

A life benefit that is a planted fraud

life claim · $49,101

The fraud W2 paid and W3 sent to a person, drawn version by version.

The claim, version by version

Every workflow's card: those tried after W3, and the two not built

Development scores flatter. Only fresh claims count.

W3 was frozen and run once on a thousand claims it had never seen, drawn from a seed chosen only when the run began. It failed: it had chosen its certainty cut-off on the same claims it was scored on. The corrected version chose it on claims its models never saw, passed, then passed a larger and stricter test.

Accuracy on fresh claims against the 99.5% bar. First attempt: failed; Second attempt: passed; Its confirmation: passed. The figures are in the text below.

Each line is the honest range of the accuracy on claims settled with no person; the dot is the measured figure, filled when the attempt passed and open when it failed.

First version

+$5,272 per 1,000 claims in development −$6,363 on fresh claims

Corrected version

+$5,939 in development +$6,790 on fresh claims, then +$12,613 on its confirmation

Dollars at the staff costs each version was frozen with.

The tests, run by run

Three attempts to settle more claims all failed on fresh claims.

Each passed in development, earned its test, and failed on fresh claims. The steps they added worked as designed. What failed each test was a claim the plain W3 inside it paid with confidence, or a side effect of the new input.

W3 + a fifteenth question

Asks the decision model whether the damage could predate the incident: the one kind of claim that caused most of W3's early mistakes.

65.6% settled alone (paid and asked together) · 4 wrong Failed

Wrongly paid $5,634; net +$13,196 per 1,000 claims.

The question caught what it was aimed at. But its answer also made W3 ask other customers for information they did not need, and a wrong request counts against accuracy.

W4: ask the customer

When one piece is missing, ask the customer for it, then decide again. A reply can settle a claim only up to a cap. Set for the lab: a reply settles at most $20,000.

59.9% settled alone (paid and asked together) · 1 wrong Failed

Wrongly paid $49,999; net −$43,069 per 1,000 claims.

The step that asks the customer made no wrong payment, and caught every attempt to manipulate it, against five fixed test phrasings; subtler attacks are not yet tested. The claim that failed the test was one plain W3 paid by itself.

W4 + a payment limit

Nothing over the limit is paid without a person, whichever step proposes it. Set for the lab: a reply settles at most $20,000; no automatic payment over $20,000.

52.2% settled alone (paid and asked together) · 1 wrong Failed

Wrongly paid $2,322; net −$3,372 per 1,000 claims.

The limit stopped the large mistake, but sent correct large payments to a person, which cost more staff time than it saved. A fault in the run also left some claims unasked, so its evidence is incomplete.

its share settled alone (paid and asked together) the corrected W3's, on its own test

What none of them fixed

A few large planted frauds that the decision model misreads.

A large planted fraud near the large-claim limit was paid in testing, by the confirmed W3 too. That is why the version now running in shadow mode pays nothing automatically above a set limit. A limit on automatic payments stops such a claim but costs more than it saves, and a language model reviewing every large payment, tried in development, flagged many honest ones and still let planted frauds through. Where to draw that line is a policy decision, not something the data can choose.

127 people by hand. 95 with the new process.

Staff time turned into the number a business plans around. The saving comes from settling routine claims with no person and from sizing the human work to the claim. The quick confirmation is kept only for requests for more information, where a missed wrong file costs a delay, not a payment. The graded fraud check is counted with a standard review for every claim its light check clears, and as a check that misses no fraud, so its part of the saving is still an upper bound. Both at 100,000 claims a year.

127 people by hand
95 with the new process

Development figures: each fold's cut-off chosen on claims its models had not seen, as in the version that passed, and the payment limit. The version that passed paid 54.9% of its fresh claims with no person and asked for documents on 4.4%; on the larger confirmation, 54.8% and 4.8%. Paid and asked apart: recomputed from the stored answers and checked against the recorded result. The headline's version paid 49.4% with no person pooled over every set, and asked for documents on 4.3%.

Change the claims a year, the hours and the minutes per task

What this shows, and what it doesn't

It shows

  • A narrow decision model plus code can settle a large share of these claims with no person, above the accuracy bar, on claims it never saw.
  • Scaling the certainty needed to the money at stake is what made that hold.
  • Development scores overstate. Every design that failed looked ready first.

It doesn't show

  • That it works on real claims. The tests on fresh claims show the result holds on new claims from the same generator, judged against answers the lab itself set; nothing yet tests it on real ones.
  • That the staff savings are real. Minutes per task are benchmark assumptions, not timed.
  • That large planted frauds are caught. One was paid in testing, by the confirmed W3 too.
  • That manipulation is reliably caught. Most attempts in the fresh claims were quarantined, but not all: some reached a decision with no person (see every wrong decision), and every attempt was built from five fixed test phrasings.
  • That it decides real claims. W3 runs in shadow mode only: it records what it would decide and changes nothing.
  • That it saves money without its payment limit. Pooled over the same fresh claims, its net per thousand claims without the limit ranges from −$8,139 to +$10,728: one wrong payment of $49,999 decides it.

Questions

Was the corrected version tuned on the test it passed?

No. The fix was worked out on the development claims only. Each test draws new claims from a seed chosen when the run starts and uses them once; no test's claims are used to fit a design. One rule did follow a test's result: the payment limit, set after W4's test made a large wrong payment. That is why the pooled money figure with the limit is a retrospective analysis. A result is final for the version tested: the first failure stays on the record.

Why synthetic claims?

Because every claim then has a known right answer, so every decision can be marked right or wrong, and frauds, manipulation attempts and missing documents can be planted on purpose. The cost is realism: a pass here is where the evidence starts, not where it ends.

Why does code decide, not the model?

A model asked to decide can be talked into a decision, and cannot say how sure it should be. Asked narrow questions, its answers can be weighed, checked against fixed rules, and priced against what a mistake would cost.

What would it take to use this on real claims?

A test on real claims, with their real outcomes; human tasks timed rather than assumed; a decision on large payments; and a period in shadow mode, where the system records what it would do while people keep deciding.

Where are the raw data and the code?

Published: a summary of every version and every test on fresh claims, and the full record of eleven example claims from one run, each file byte for byte as it was exported. Every figure on this site traces to one of them. Not published: the records of the other claims, and the code, which is in a private repository. The manifest names each file and its hash.

Disclosure: the lab has no relationship with TypeSafe, the company that makes Jev. Jev was used through OpenRouter, and the lab's owner paid for every model call, Jev's and the language models', personally. The lab's owner designed, ran and judged every test, working with AI coding assistants; nothing here has yet been replicated independently.