Pharmacy demos

Can code handle routine pharmacy decisions, with a narrow model answering?

Three small demos of one design. Jev, a narrow decision model, answers typed questions about a report or request. Code decides the route. A person reviews everything code doesn’t settle, and nothing is denied automatically.

Demo, not validated · not clinical advice

Figures last updated

The three demos, as they stand

01 · Vaccine reports80 of 100serious reports sent to a safety case at the built line of 0.5 (range 71.1 to 86.7%); 7 of 200 others flagged.
02 · Prior authorisation72 of 178complete requests approved in the last design (range 33.5 to 47.8%); wrongly approved, none of 72 (up to 5.1%). As first built: 0 of 178.
03 · Refills13 of 29automatic approvals wrong as built (44.8%, range 28.4 to 62.5%), all overdue refills. Re-routed afterwards on the same answers, with code checking the dates (right by construction here): none wrong, but only 20 of 110 approved. With the manipulation answer also left out of approval: 104 of 110 approved, and 1 of the 9 manipulation attempts approved (1 of the 6 reworded ones).

Every figure on this page comes from one run of each demo on a few hundred cases, and some settings were chosen on those same cases. None of it is validated, and none of it is clinical advice.

01 · Vaccine-report triage

Which adverse-event reports should a person look at first for the fifteen-day clock?

Demo, not validated · not clinical advice

A drug maker must send serious, unexpected adverse events to the regulator within fifteen days. Someone reads every incoming report to decide. The demo asks whether Jev can flag the serious ones in real public reports (VAERS), so a person starts the clock sooner.

How it works

PeopleJev, the decision modelCodeWhere it ends upA report arrivesFour narrow answersRoute by a lineCausality and the reportSafety casethe fifteen-day flagRoutineread in the normal queue

The result, at the built line

Serious reports sent to a safety case80 of 100range 71.1 to 86.7%, as coded by the reporter
Of them, with no label in the narrative56 of 76range 62.8 to 82.3%: what Jev does without the sender’s own verdict
Others flagged as serious7 of 200range 1.7 to 7.0%, a person’s time
Jev, per reportabout $0.00005$0.0151 for 300 reports

Settled with no person: none. Every report still reaches a person, in the safety queue or the routine one; the demo only sorts which a person reads first. Any gain would be speed, starting the clock sooner, and that is not measured.

One mark per report the reporter coded serious: filled if sent to a safety case, open if it went to the routine queue (missed).

Moving the line

LineSerious caughtOthers flagged
0.5 (built)80 of 100 (80.0%, range 71.1 to 86.7%)7 of 200 (3.5%, range 1.7 to 7.0%)
0.483 of 100 (83.0%, range 74.5 to 89.1%)13 of 200 (6.5%, range 3.8 to 10.8%)
0.390 of 100 (90.0%, range 82.6 to 94.5%)21 of 200 (10.5%, range 7.0 to 15.5%)
0.2 (recommended)95 of 100 (95.0%, range 88.8 to 97.8%)36 of 200 (18.0%, range 13.3 to 23.9%)

The lab’s recommendation for safety triage is 0.2: a missed report risks the clock, an extra one costs minutes. It is a recommendation from a small demo, not a validated setting, and it was chosen on these same reports, so its rates here are optimistic.

Reweighted to the pool’s mix (7.9% coded serious), an estimate: at 0.5, about 66% of flagged reports would be serious and about 9.5% of all reports flagged; at 0.2, about 31% and 24.1%.

How far reading the label alone gets you

Some narratives state the sender’s own seriousness verdict: 83 of 300. Reading that label alone would catch 21 of the 100 serious reports; Jev caught 80, including 21 of those 21. The label is no answer key either: it disagrees with the reporter’s coding on 20 of the 83. The other 217 narratives carry no label (76 of them coded serious), and that is where all 20 misses fall.

What this shows

  • Jev can read messy public narratives and flag most coded-serious reports, at a cost of cents.
  • The misses sit where no label helps; the most common coding among them is disability (10 of 20).

What it doesn’t

  • VAERS reports are voluntary and unverified; they don’t show a vaccine caused anything.
  • “Serious” is what the reporter coded, and the sample over-represents it: 100 of 300 here, against 7.9% in the pool.
  • The clock needs serious and unexpected; Jev isn’t asked about expectedness, so the route is a triage flag.
  • Some reports (12) shared an ID with another VAERS report, and Jev saw both reports’ vaccines.
  • One run, one sample: no held-out test, no repeat.
  • Time and money saved: not measured. A saving needs the minutes a person spends on each case and what that time costs, and these demos measure neither.

02 · Drug prior authorisation

Can code approve a prior-authorisation request when every criterion is met?

Demo, not validated · not clinical advice

A pharmacist or payer reviewer reads each request against the drug’s criteria. Most meet them. The demo asks whether code can approve those on its own, from Jev’s answers, and send everything else to a person. It never denies.

How it works

PeopleJev, the decision modelCodeWhere it ends upFacts and a noteSeven narrow answersHard rules, then a lineEverything not approvedApprovedevery criterion clearly metAsked for an itemback to the prescriberTo a pharmacistanything uncertain

Every run, failures included

As first built0 of 178Plus the request date0 of 178Plus the required items0 of 178A gentler rule78 of 178, 2 wrongLast design72 of 178noneAll complete requests

Hatched: that run also approved requests it should not have. Each run is scored under its own rule.

RunWhat changedScored underComplete approvedWrongly approvedApproved with no personSent back with no personJev; notes (set-up)
As first builtThe design as registered: the weakest of seven answers at 0.9, cost scaled.every answer at 0.9 or more, cost scaled in full0 of 178 (0.0%, range 0.0 to 2.1%)none approved0 of 289129 of 289$0.0094; $0.2996
Plus the request dateThe first run’s requests and notes, with only the request date added.every answer at 0.9 or more, cost scaled in full0 of 178 (0.0%, range 0.0 to 2.1%)none approved0 of 289128 of 289$0.0096; notes reused
Plus the required itemsThe first run’s requests and notes, with only each drug’s required items added.every answer at 0.9 or more, cost scaled in full0 of 178 (0.0%, range 0.0 to 2.1%)none approved0 of 28973 of 289$0.0101; notes reused
A gentler ruleDate and required items, every answer at 0.7 on a gentler cost scale, new requests.every answer at 0.7 or more, cost scaled gently78 of 178 (43.8%, range 36.7 to 51.2%)2 of 80 (range 0.7 to 8.7%)80 of 29188 of 291$0.0110; $0.2929
Last designThe gentler rule with contradictions that truly rule out the diagnosis, and a screening question.every answer at 0.7 or more, cost scaled gently72 of 178 (40.4%, range 33.5 to 47.8%)none of 72 (up to 5.1%)72 of 29184 of 291$0.0119; $0.2977

Work taken off a person, and what the model cost

Approved with no person, last design72 of 291answered requests (24.7%, range 20.1 to 30.0%)
Sent back for an item, with no person84 of 291(28.9%, range 24.0 to 34.3%): no reviewer at that step, but the request usually comes back, and the demo does not follow it
Jev, per requestabout $0.00004the model’s cost to answer one request
Writing the notes, per requestabout $0.0010a cost of the synthetic set-up only: real requests arrive with their notes

What made the difference: the first run’s notes, one change at a time

On the first run’s requests and notesScored underComplete approvedWrongly approved
The gentler rule aloneevery answer at 0.7 or more, cost scaled gently40 of 178 (22.5%, range 17.0 to 29.1%)none of 40 (up to 8.8%)
Plus the request dateevery answer at 0.7 or more, cost scaled gently43 of 178 (24.2%, range 18.5 to 30.9%)none of 43 (up to 8.2%)
Plus each drug’s required itemsevery answer at 0.7 or more, cost scaled gently76 of 178 (42.7%, range 35.7 to 50.0%)3 of 79 (range 1.3 to 10.6%)

On the same notes, the gentler rule alone approves 40 of 178 complete requests, and adding the required-items list takes it to 76; the list removes the false “item missing” answers. The request date alone changes little: 43 against 40 under the gentler rule, and nothing under the other two. Drug-level differences from the gentler rule’s own run (on new requests, 78 of 178) are not separated from its new seed and notes.

In the last design, 15 of 15 contradictions went to a pharmacist. Adalimumab, the dearest drug, still got none of its 24 complete requests approved: its cost raises the line above what Jev’s answers reached. The wrong approvals in the gentler rule’s own run, and under it on the first run’s notes, came from a fault in the answer key, a “contradiction” that wasn’t one; the last design’s generator fixes it.

What this shows

  • In the last design, code approved about four in ten complete requests (72 of 178, range 33.5 to 47.8%), with none of 72 (up to 5.1%) wrongly approved. As first built, it approved 0 of 178.
  • Which change mattered, where the runs isolate it: the required-items list and a gentler rule, together.

What it doesn’t

  • The requests are synthetic and the notes written by a language model, so this shows how the workflow behaves, not real-world accuracy.
  • The answer key is the generator’s own facts, and it was wrong once.
  • The later designs were built after seeing the earlier runs. The headline is the last design, shown beside the design as first built.
  • Manipulation checks are scoped to fixed test phrasings, and only 6 were asked in the last run.
  • The predictions were written before each run; some were informed by earlier runs, and each is marked in the research session’s record.
  • Time and money saved: not measured. A saving needs the minutes a person spends on each case and what that time costs, and these demos measure neither.

03 · Refills under protocol

Can code refill a prescription under a clinic’s protocol?

Demo, not validated · not clinical advice

Clinics renew routine prescriptions under standing orders: labs current, a recent visit, no new problem. The demo asks whether code can approve those from Jev’s reading of the chart note, while its own rules, from the fill history alone, send controlled drugs, early refills and high-alert drugs to a pharmacist.

How it works

PeopleJev, the decision modelCodeWhere it ends upFill history checkedSeven narrow answersRoute by risk lineEverything not approvedApprovedwithin protocolLabs or a visitback to the prescriberTo a pharmacistcode’s rules or doubt

Three designs, the same answers

Each design routes the same requests, the same notes and the same Jev answers; only code’s rules change, and nothing was asked again. As built, Jev judged the lab and visit dates. The second design makes two changes: code reads the dates from structured records, and Jev’s two date answers leave the approval decision. The third also leaves Jev’s manipulation answer out of the approval decision, while still quarantining a clear attempt. Here the dates code reads are the generator’s own facts, the ones the answer key is built from, so the date check is right by construction; real records can be missing, late or wrong, and the gain would be smaller.

The trade-off. As built: 13 wrong approvals, all overdue refills. Code checks the dates: none wrong, and 45 of 45 requests needing labs or a visit sent back, but only 20 of 110 approved. Manipulation answer out of approval: 104 of 110 approved, and 1 of the 9 manipulation attempts approved (1 of the 6 reworded ones).

DesignWhat code doesApproved with no personWrongly approvedLabs or a visit, asked forMatches the facts
As builtJev judges the dates16 of 110 (14.5%, range 9.2 to 22.3%)13 of 29 (range 28.4 to 62.5%)0 of 45158 of 300
Date check alone (a reading)code’s date check, with Jev’s date answers still in the approval decision; computed here from the stored answers, not a design the lab’s code ran16 of 110 (14.5%, range 9.2 to 22.3%)none of 16 (up to 19.4%)45 of 45203 of 300
Code checks the datesreads the dates from structured records, and drops Jev’s two date answers from the approval decision20 of 110 (18.2%, range 12.1 to 26.4%)none of 20 (up to 16.1%)45 of 45207 of 300
Manipulation answer out of approvalas the second, and the manipulation answer only quarantines104 of 110 (94.5%, range 88.6 to 97.5%)1 of 105 (range 0.2 to 5.2%)45 of 45291 of 300

Where each gain comes from, computed from the stored answers: the date check alone removes every wrong approval (0 of 16) and leaves approvals at 16 of 110; dropping Jev’s two date answers from the approval decision is what takes them to 20.

Jev’s manipulation answer does not separate clean notes from reworded attempts at the low end. Under the third design, approved requests scored up to 0.49 on it; the reworded attempt it approved scored 0.36, and 42 clean requests it approved scored at least as much. No line on that answer alone could have stopped the attempt without also holding back those clean requests.

Both re-routes were chosen after seeing earlier results on these same requests: the second after the first run, the third after the second. Their figures are optimistic, and neither has been tested on new requests.

As built, in detail

As built, code isn’t safe to approve refills on its own here: of its 29 automatic approvals, 13 were wrong (44.8%, range 28.4 to 62.5%), every one an overdue refill.

Approved with no person16 of 110requests the facts say should be approved (14.5%, range 9.2 to 22.3%). In all, 29 of 300 requests were approved with no person, 13 of them wrongly.
Sent back for labs or a visit, with no person0 of 300No overdue request was sent back: of the 53 planted with a lab or visit overdue, 40 went to a pharmacist and 13 were approved.
Automatic approvals that were wrong13 of 29(44.8%, range 28.4 to 62.5%): lab overdue 7, visit overdue 6, every one an overdue refill
Jev, per requestabout $0.00003$0.0096 for 300. Writing the notes ($0.2569) is a cost of the synthetic set-up only: real requests arrive with their notes.
One mark per request the facts say should be approved: filled if code approved it, open if it went to a person.

The main finding: Jev didn’t judge dates

Jev answered “labs current” for 23 of the 23 requests planted with an overdue lab, and “visit current” for 30 of the 30 with an overdue visit. So code never asked for labs or a visit, and all 13 wrong approvals were overdue refills. A real deployment would check those dates in code from structured records, as the fill rules already do.

Approvals were rare for a second reason: of the 91 requests that should have been approved and fell short of the line on Jev’s answers, the manipulation answer was the weakest in 83. It sits well above zero on ordinary notes: its median on requests within protocol was 0.34.

What worked

Lab out of range, to a pharmacist22 of 22range 85.1 to 100.0%
Dose change or side effect, to a pharmacist24 of 24range 86.2 to 100.0%
Contradictions, to a pharmacist12 of 12range 75.8 to 100.0%
Manipulation attempts approved0 of 9Asked verbatim: 3 of 3 quarantined. Reworded: 4 of 6 quarantined, 2 to a pharmacist.

Every request, by what was planted

PlantedRequestsApprovedLabs or a visitTo a pharmacistQuarantinedMatches the facts
Within protocol165160146371
Lab overdue23701601
Visit overdue30602407
Lab out of range220022022
Dose change or side effect240024024
Early refill150014114
Contradiction120012012
Manipulation attempt900277

All 300 requests were answered. In the first phase the note check refused 22 notes: 16 as too short, under a floor since lowered, and 6 manipulation attempts the writer had reworded in its own voice, which the check did not yet recognise. All 22 passed the corrected check in phase two and were asked; the reworded attempts are the 6 counted under manipulation.

89 requests went to a pharmacist on code’s own fill rules, whatever Jev answered: controlled drugs, early refills and high-alert drugs. Jev never sees the fill data.

What this shows

  • As built, it isn’t safe to approve on its own: 13 of its 29 automatic approvals were wrong (44.8%), every one an overdue refill. The date finding explains why.
  • Re-routed after the fact on the same requests, moving the date check into code removed the wrong approvals (0 of 20; all the as-built ones were overdue refills) and sent back 45 of 45 requests needing labs or a visit. Here that check is right by construction, and the approval count reflects both of the second design’s changes together.
  • Code’s fill rules and Jev’s clinical flags together kept every planted problem except an overdue date from being approved.
  • As built, no manipulation attempt was approved (0 of 9); with the manipulation answer left out of approval, 1 of 9 was.
  • Code approved about one in seven requests that should be approved (16 of 110), so most still reach a person.

What it doesn’t

  • The requests are synthetic and the notes written by a language model, so this shows how the workflow behaves, not real-world accuracy.
  • The protocols are illustrative, not a real clinic’s standing orders.
  • The second and third designs were chosen after seeing earlier results on the same requests, so their figures are optimistic, and neither has been tested on new requests.
  • The dates the second and third designs read are the generator’s own facts, the ones the answer key is built from, so their date check is right by construction; real records can be missing, late or wrong.
  • 4 ordinary requests were quarantined as manipulation, all on notes in the shortest quarter.
  • The fill rules are the fraud guard’s first piece only: they use the fill history, not identity, prescriber or pharmacy checks.
  • Manipulation checks are scoped to fixed test phrasings and rewordings of them, and only 9 were asked.
  • One run on 300 requests. The research session scored this demo’s predictions: 77 and 78 missed, and 79 to 85 hit, by their numbers in its record.
  • Time and money saved: not measured. A saving needs the minutes a person spends on each case and what that time costs, and these demos measure neither.

04 · What a real deployment would need

From three demos to a pharmacy’s real queue

05 · How far to trust these numbers

One run each, small samples

Every figure is from a single run on a few hundred cases, and every rate carries its ninety-five per cent range (a Wilson interval). Where a setting was chosen after seeing a run, the page says so and its figures are optimistic. The predictions were written before each run and scored by the lab’s research session; some were informed by earlier runs or by sample-level information, and each is marked in its record.

The research session’s score across the pharmacy predictions: 20 hits of 27 scored, all marked informed in the record.

DemoRunsJevWriting the notes, a synthetic set-up cost
Vaccine reports1$0.0151none: the reports are public
Prior authorisation5$0.0520$0.8902
Refills1$0.0096$0.2569

Where each figure comes from

Each run’s full export is held in the lab’s private repository, and this page is computed from a frozen extract of it with no report or request text. Every route on the page is re-derived from Jev’s answers and the run’s own rule constants when the page is built, and the build fails if one differs. The synthetic notes were written by a language model, named in the table; the VAERS reports are public.

DemoRunModelsExportRules read from
Vaccine reports8b7e6a5e-03b4-498a-9b89-587f68231b25~typesafe/jev-latest, answered by typesafe/jev-1.13-20260917 (all 300)no notes: the reports are publicsha256 6a9c236526f7bf8dc52cb7102063c7e5c872f3c11f33df768907bcdcaf800a2fSERIOUS_LINE, SAYS_SERIOUS and SAYS_NOT_SERIOUS in apps/api/src/probes/vaers.ts at 968a8669d5ad51a2a4f32a60020e0d9e19270876
Prior authorisationAs first built538767d4-3f54-4863-9fc5-3840674141e0~typesafe/jev-latest, answered by typesafe/jev-1.13-20260917 (all rows)notes written by anthropic/claude-haiku-4.5sha256 a52567513c8e04009b514b927e0e0dc974041b44b9c16113418a89af31a644b0PA_LINE and COST_REFERENCE_USD in apps/api/src/probes/pa.ts at 98773469d40df4c55f7b47649b3e98eea020af86
Prior authorisationPlus the request date0ba06f70-c1a8-480b-a2d8-4d322e239800~typesafe/jev-latest, answered by typesafe/jev-1.13-20260917 (all rows), question set PA1notes written by anthropic/claude-haiku-4.5sha256 a79663130bc5554206ab6b4c7c6f16c2f936f0bab27b402c5709fcfb3590a2a9the run's export (design 3, as DESIGNS in apps/api/src/probes/pa.ts at 311da76a071df69e164303a8cdf6b2b63795f4db) and COST_REFERENCE_USD there
Prior authorisationPlus the required itemsd97049c3-7336-4805-9b6a-da8ca994d4a1~typesafe/jev-latest, answered by typesafe/jev-1.13-20260917 (all rows), question set PA1notes written by anthropic/claude-haiku-4.5sha256 d7f82f99b1d5ef3c5ed68c2d7b4a845144477b2bbd5619fbc1e90cbd49f80faathe run's export (design 5, as DESIGNS in apps/api/src/probes/pa.ts at ac6b459a6fc4df38799822e49ebeca48c72c3f6d) and COST_REFERENCE_USD there
Prior authorisationA gentler rule58e1ce6a-1196-47e4-bad3-81dcc072b279~typesafe/jev-latest, answered by typesafe/jev-1.13-20260917 (all rows), question set PA2notes written by anthropic/claude-haiku-4.5sha256 81cfd3c2ea94db84dbdc86bd6d71f9b059cd0c8a256360683dd3147c156acebathe run's export (design 2, as DESIGNS in apps/api/src/probes/pa.ts at bb16334b33ac44ed9bdb838e0ea77cf20da9daba) and COST_REFERENCE_USD there
Prior authorisationLast design6f0414c7-d2b4-41ec-8be8-503b385b46e0~typesafe/jev-latest, answered by typesafe/jev-1.13-20260917 (all rows), question set PA3notes written by anthropic/claude-haiku-4.5sha256 f4b791a54ab780404eef945ce67a960818e63fa0150c24335a5ffdf20c177a53the run's export (design 4, as DESIGNS in apps/api/src/probes/pa.ts at ac6b459a6fc4df38799822e49ebeca48c72c3f6d) and COST_REFERENCE_USD there
Refills85750433-99de-4318-aed8-ecfac36e5a7e~typesafe/jev-latest, answered by typesafe/jev-1.13-20260917 (all rows), question set RF1notes written by anthropic/claude-haiku-4.5sha256 0ea867783b2c3ee4a9eb480de1691108ee5fc412fef8737ff0571eda092e15e6RF_LINES, EARLY_SHARE, REFILL_LIMIT, REQUEST_DATE and the formulary in apps/api/src/probes/rf.ts at 54a5a4c0ed57b074130471100f48497bc9d9b4dd

Disclosure: the lab has no relationship with TypeSafe, the company that makes Jev. Jev was used through OpenRouter, and the lab’s owner paid for every model call, Jev’s and the language models’, personally. The lab’s owner designed, ran and judged every demo, working with AI coding assistants; nothing here has been replicated independently.