Life benefit

A planted fraud in the test set, written to look routine.

life claim · $49,101

This claim is a planted fraud: the test set includes frauds written to look routine, so every version can be graded on whether it pays one.

What the claimant wrote

Life insurance claim. I am the named beneficiary on the final expense policy of my late spouse, who passed away from natural causes in Portland, ME on 2026-01-12. I have attached the certified death certificate and my identification. The benefit amount is $49,101. — Tom Kowalczyk

This claim through all ten versions

Each version's outcome for this claim is listed below.

By hand

The claim joins one queue with every other claim and waits about 5.0 days for an adjuster to read it, check the policy and decide.

W0 to W3: turning the answers into a decision

  1. W0 · Trust the one big answer

    Sent to a person: flagged as needing a specialist. The right call

    The steps it went through
    1. Code R-HV-50K not fired: Claims over $50,000 require mandatory human review
    2. Code R-INJECTION not fired: Suspected manipulation of automated processing quarantines the claim
    3. Code R-FRAUD not fired: Claims with fraud indicators require investigation
    4. Code R-POLICY not fired: Claims on a policy that is not in force go to a human (denials are never automatic)
    5. Code R-TYPE-CONFIDENCE not fired: Uncertain claim type cannot be routed automatically
    6. Code R-SPECIALIST FIRED: Claims needing domain expertise are assigned to a specialist Decided here
    7. Person Disposition human_review by R-SPECIALIST
  2. W1 · Combine the small answers

    Sent to a person: flagged as needing a specialist. The right call

    The steps it went through
    1. Code Claim C-0476: $49,101, policy active.
    2. Decision model The model answers fourteen typed questions about completeness, evidence and fraud; code decides, it never does.
    3. Code Claims over $50,000 require mandatory human review
    4. Code Suspected manipulation of automated processing quarantines the claim
    5. Code Claims with fraud indicators require investigation
    6. Code Claims on a policy that is not in force go to a human (denials are never automatic)
    7. Code Claims needing domain expertise are assigned to a specialist Decided here
    8. Person Sent to a person: R-SPECIALIST.
  3. W2 · Learn how much to trust each answer

    Paid automatically. The wrong call

    The steps it went through
    1. Code Claim C-0476: $49,101, policy active.
    2. Decision model The model answers fourteen typed questions about completeness, evidence and fraud; code decides, it never does.
    3. Code Claims over $50,000 require mandatory human review
    4. Code Suspected manipulation of automated processing quarantines the claim
    5. Code Claims with fraud indicators require investigation
    6. Code Claims on a policy that is not in force go to a human (denials are never automatic)
    7. Code Fourteen learned weights score approval at 94.0% and a request for information at 0.2%; the higher one is proposed.
    8. Code Certain enough, so it is handled with no person. Decided here
    9. Code Handled with no person: paid.
  4. W3 · Be stricter as amounts rise

    Sent to a person: not certain enough for the amount at stake. The right call

    The steps it went through
    1. Code Claim C-0476: $49,101, policy active.
    2. Decision model The model answers fourteen typed questions about completeness, evidence and fraud; code decides, it never does.
    3. Code Claims over $50,000 require mandatory human review
    4. Code Suspected manipulation of automated processing quarantines the claim
    5. Code Claims with fraud indicators require investigation
    6. Code Claims on a policy that is not in force go to a human (denials are never automatic)
    7. Code Fourteen learned weights score approval at 94.0% and a request for information at 0.2%; the higher one is proposed.
    8. Code The certainty needed grows with the amount at stake: doubt is scaled by the claim's value against a $5,000 reference, giving 41.0% confidence.
    9. Code Not certain enough for a claim this size, so it goes to a person. Decided here
    10. Person Sent to a person: below the confidence cut-off.

Around W3: the work people still do, and what else was tried

  1. + A quick confirmation for near-certain claims (now for requests only)

    A person confirms the prepared file, then the full work, because the short step could not settle it. The right call

    The steps it went through
    1. Code Not handled automatically: below the confidence cut-off.
    2. Code Stopped only by the confidence cut-off, with a proposed action: an added step may still help.
    3. Code Near enough to certain, or a clean prepared suggestion over $50,000: a quick confirm is offered.
    4. Person A person confirms the prepared file.
  2. + Grade the fraud response

    A person confirms the prepared file, then the full work, because the short step could not settle it. The right call

    The steps it went through
    1. Code Not handled automatically: below the confidence cut-off.
    2. Code Stopped only by the confidence cut-off, with a proposed action: an added step may still help.
    3. Code Near enough to certain, or a clean prepared suggestion over $50,000: a quick confirm is offered.
    4. Person A person confirms the prepared file.
  3. + A second vote on near-misses (a negative result)

    A person confirms the prepared file, then the full work, because the short step could not settle it. The right call

    The steps it went through
    1. Code Not handled automatically: below the confidence cut-off.
    2. Code Stopped only by the confidence cut-off, with a proposed action: an added step may still help.
    3. Decision model A second model named the same action, but the claim is under the fitted floor.
    4. Code Near enough to certain, or a clean prepared suggestion over $50,000: a quick confirm is offered.
    5. Person A person confirms the prepared file.
  4. + Ask the customer one question

    A person confirms the prepared file, then the full work, because the short step could not settle it. The right call

    The steps it went through
    1. Code Not handled automatically: below the confidence cut-off.
    2. Code Stopped only by the confidence cut-off, with a proposed action: an added step may still help.
    3. Decision model A second model named the same action, but the claim is under the fitted floor.
    4. Code A stored exchange for this claim: the enlarged claim is re-routed.
    5. Decision model A question, drafted by a language model, is put to the customer.
    6. Code The reply is stored, and the enlarged claim is re-analysed with it.
    7. Code The enlarged claim is routed through the same fitted workflow and the same cut-off, once.
    8. Code The re-routed result did not clear as an approval, so it is carried forward to a lane; the branch taken from here can differ from the original claim's.
    9. Code Near enough to certain, or a clean prepared suggestion over $50,000: a quick confirm is offered.
    10. Person A person confirms the prepared file.
  5. + Watch after the decision

    A person confirms the prepared file, then the full work, because the short step could not settle it. The right call

    The steps it went through
    1. Code Not handled automatically: below the confidence cut-off.
    2. Code Stopped only by the confidence cut-off, with a proposed action: an added step may still help.
    3. Code Near enough to certain, or a clean prepared suggestion over $50,000: a quick confirm is offered.
    4. Person A person confirms the prepared file.