Pharmacy demos
Can code handle routine pharmacy decisions, with a narrow model answering?
Three small demos of one design. Jev, a narrow decision model, answers typed questions about a report or request. Code decides the route. A person reviews everything code doesn’t settle, and nothing is denied automatically.
Demo, not validated · not clinical advice
Figures last updated
The three demos, as they stand
Every figure on this page comes from one run of each demo on a few hundred cases, and some settings were chosen on those same cases. None of it is validated, and none of it is clinical advice.
01 · Vaccine-report triage
Which adverse-event reports should a person look at first for the fifteen-day clock?
Demo, not validated · not clinical advice
A drug maker must send serious, unexpected adverse events to the regulator within fifteen days. Someone reads every incoming report to decide. The demo asks whether Jev can flag the serious ones in real public reports (VAERS), so a person starts the clock sooner.
How it works
The result, at the built line
Settled with no person: none. Every report still reaches a person, in the safety queue or the routine one; the demo only sorts which a person reads first. Any gain would be speed, starting the clock sooner, and that is not measured.
Moving the line
| Line | Serious caught | Others flagged |
|---|---|---|
| 0.5 (built) | 80 of 100 (80.0%, range 71.1 to 86.7%) | 7 of 200 (3.5%, range 1.7 to 7.0%) |
| 0.4 | 83 of 100 (83.0%, range 74.5 to 89.1%) | 13 of 200 (6.5%, range 3.8 to 10.8%) |
| 0.3 | 90 of 100 (90.0%, range 82.6 to 94.5%) | 21 of 200 (10.5%, range 7.0 to 15.5%) |
| 0.2 (recommended) | 95 of 100 (95.0%, range 88.8 to 97.8%) | 36 of 200 (18.0%, range 13.3 to 23.9%) |
The lab’s recommendation for safety triage is 0.2: a missed report risks the clock, an extra one costs minutes. It is a recommendation from a small demo, not a validated setting, and it was chosen on these same reports, so its rates here are optimistic.
Reweighted to the pool’s mix (7.9% coded serious), an estimate: at 0.5, about 66% of flagged reports would be serious and about 9.5% of all reports flagged; at 0.2, about 31% and 24.1%.
How far reading the label alone gets you
Some narratives state the sender’s own seriousness verdict: 83 of 300. Reading that label alone would catch 21 of the 100 serious reports; Jev caught 80, including 21 of those 21. The label is no answer key either: it disagrees with the reporter’s coding on 20 of the 83. The other 217 narratives carry no label (76 of them coded serious), and that is where all 20 misses fall.
What this shows
- Jev can read messy public narratives and flag most coded-serious reports, at a cost of cents.
- The misses sit where no label helps; the most common coding among them is disability (10 of 20).
What it doesn’t
- VAERS reports are voluntary and unverified; they don’t show a vaccine caused anything.
- “Serious” is what the reporter coded, and the sample over-represents it: 100 of 300 here, against 7.9% in the pool.
- The clock needs serious and unexpected; Jev isn’t asked about expectedness, so the route is a triage flag.
- Some reports (12) shared an ID with another VAERS report, and Jev saw both reports’ vaccines.
- One run, one sample: no held-out test, no repeat.
- Time and money saved: not measured. A saving needs the minutes a person spends on each case and what that time costs, and these demos measure neither.
02 · Drug prior authorisation
Can code approve a prior-authorisation request when every criterion is met?
Demo, not validated · not clinical advice
A pharmacist or payer reviewer reads each request against the drug’s criteria. Most meet them. The demo asks whether code can approve those on its own, from Jev’s answers, and send everything else to a person. It never denies.
How it works
Every run, failures included
Hatched: that run also approved requests it should not have. Each run is scored under its own rule.
| Run | What changed | Scored under | Complete approved | Wrongly approved | Approved with no person | Sent back with no person | Jev; notes (set-up) |
|---|---|---|---|---|---|---|---|
| As first built | The design as registered: the weakest of seven answers at 0.9, cost scaled. | every answer at 0.9 or more, cost scaled in full | 0 of 178 (0.0%, range 0.0 to 2.1%) | none approved | 0 of 289 | 129 of 289 | $0.0094; $0.2996 |
| Plus the request date | The first run’s requests and notes, with only the request date added. | every answer at 0.9 or more, cost scaled in full | 0 of 178 (0.0%, range 0.0 to 2.1%) | none approved | 0 of 289 | 128 of 289 | $0.0096; notes reused |
| Plus the required items | The first run’s requests and notes, with only each drug’s required items added. | every answer at 0.9 or more, cost scaled in full | 0 of 178 (0.0%, range 0.0 to 2.1%) | none approved | 0 of 289 | 73 of 289 | $0.0101; notes reused |
| A gentler rule | Date and required items, every answer at 0.7 on a gentler cost scale, new requests. | every answer at 0.7 or more, cost scaled gently | 78 of 178 (43.8%, range 36.7 to 51.2%) | 2 of 80 (range 0.7 to 8.7%) | 80 of 291 | 88 of 291 | $0.0110; $0.2929 |
| Last design | The gentler rule with contradictions that truly rule out the diagnosis, and a screening question. | every answer at 0.7 or more, cost scaled gently | 72 of 178 (40.4%, range 33.5 to 47.8%) | none of 72 (up to 5.1%) | 72 of 291 | 84 of 291 | $0.0119; $0.2977 |
Work taken off a person, and what the model cost
What made the difference: the first run’s notes, one change at a time
| On the first run’s requests and notes | Scored under | Complete approved | Wrongly approved |
|---|---|---|---|
| The gentler rule alone | every answer at 0.7 or more, cost scaled gently | 40 of 178 (22.5%, range 17.0 to 29.1%) | none of 40 (up to 8.8%) |
| Plus the request date | every answer at 0.7 or more, cost scaled gently | 43 of 178 (24.2%, range 18.5 to 30.9%) | none of 43 (up to 8.2%) |
| Plus each drug’s required items | every answer at 0.7 or more, cost scaled gently | 76 of 178 (42.7%, range 35.7 to 50.0%) | 3 of 79 (range 1.3 to 10.6%) |
On the same notes, the gentler rule alone approves 40 of 178 complete requests, and adding the required-items list takes it to 76; the list removes the false “item missing” answers. The request date alone changes little: 43 against 40 under the gentler rule, and nothing under the other two. Drug-level differences from the gentler rule’s own run (on new requests, 78 of 178) are not separated from its new seed and notes.
In the last design, 15 of 15 contradictions went to a pharmacist. Adalimumab, the dearest drug, still got none of its 24 complete requests approved: its cost raises the line above what Jev’s answers reached. The wrong approvals in the gentler rule’s own run, and under it on the first run’s notes, came from a fault in the answer key, a “contradiction” that wasn’t one; the last design’s generator fixes it.
What this shows
- In the last design, code approved about four in ten complete requests (72 of 178, range 33.5 to 47.8%), with none of 72 (up to 5.1%) wrongly approved. As first built, it approved 0 of 178.
- Which change mattered, where the runs isolate it: the required-items list and a gentler rule, together.
What it doesn’t
- The requests are synthetic and the notes written by a language model, so this shows how the workflow behaves, not real-world accuracy.
- The answer key is the generator’s own facts, and it was wrong once.
- The later designs were built after seeing the earlier runs. The headline is the last design, shown beside the design as first built.
- Manipulation checks are scoped to fixed test phrasings, and only 6 were asked in the last run.
- The predictions were written before each run; some were informed by earlier runs, and each is marked in the research session’s record.
- Time and money saved: not measured. A saving needs the minutes a person spends on each case and what that time costs, and these demos measure neither.
03 · Refills under protocol
Can code refill a prescription under a clinic’s protocol?
Demo, not validated · not clinical advice
Clinics renew routine prescriptions under standing orders: labs current, a recent visit, no new problem. The demo asks whether code can approve those from Jev’s reading of the chart note, while its own rules, from the fill history alone, send controlled drugs, early refills and high-alert drugs to a pharmacist.
How it works
Three designs, the same answers
Each design routes the same requests, the same notes and the same Jev answers; only code’s rules change, and nothing was asked again. As built, Jev judged the lab and visit dates. The second design makes two changes: code reads the dates from structured records, and Jev’s two date answers leave the approval decision. The third also leaves Jev’s manipulation answer out of the approval decision, while still quarantining a clear attempt. Here the dates code reads are the generator’s own facts, the ones the answer key is built from, so the date check is right by construction; real records can be missing, late or wrong, and the gain would be smaller.
The trade-off. As built: 13 wrong approvals, all overdue refills. Code checks the dates: none wrong, and 45 of 45 requests needing labs or a visit sent back, but only 20 of 110 approved. Manipulation answer out of approval: 104 of 110 approved, and 1 of the 9 manipulation attempts approved (1 of the 6 reworded ones).
| Design | What code does | Approved with no person | Wrongly approved | Labs or a visit, asked for | Matches the facts |
|---|---|---|---|---|---|
| As built | Jev judges the dates | 16 of 110 (14.5%, range 9.2 to 22.3%) | 13 of 29 (range 28.4 to 62.5%) | 0 of 45 | 158 of 300 |
| Date check alone (a reading) | code’s date check, with Jev’s date answers still in the approval decision; computed here from the stored answers, not a design the lab’s code ran | 16 of 110 (14.5%, range 9.2 to 22.3%) | none of 16 (up to 19.4%) | 45 of 45 | 203 of 300 |
| Code checks the dates | reads the dates from structured records, and drops Jev’s two date answers from the approval decision | 20 of 110 (18.2%, range 12.1 to 26.4%) | none of 20 (up to 16.1%) | 45 of 45 | 207 of 300 |
| Manipulation answer out of approval | as the second, and the manipulation answer only quarantines | 104 of 110 (94.5%, range 88.6 to 97.5%) | 1 of 105 (range 0.2 to 5.2%) | 45 of 45 | 291 of 300 |
Where each gain comes from, computed from the stored answers: the date check alone removes every wrong approval (0 of 16) and leaves approvals at 16 of 110; dropping Jev’s two date answers from the approval decision is what takes them to 20.
Jev’s manipulation answer does not separate clean notes from reworded attempts at the low end. Under the third design, approved requests scored up to 0.49 on it; the reworded attempt it approved scored 0.36, and 42 clean requests it approved scored at least as much. No line on that answer alone could have stopped the attempt without also holding back those clean requests.
Both re-routes were chosen after seeing earlier results on these same requests: the second after the first run, the third after the second. Their figures are optimistic, and neither has been tested on new requests.
As built, in detail
As built, code isn’t safe to approve refills on its own here: of its 29 automatic approvals, 13 were wrong (44.8%, range 28.4 to 62.5%), every one an overdue refill.
The main finding: Jev didn’t judge dates
Jev answered “labs current” for 23 of the 23 requests planted with an overdue lab, and “visit current” for 30 of the 30 with an overdue visit. So code never asked for labs or a visit, and all 13 wrong approvals were overdue refills. A real deployment would check those dates in code from structured records, as the fill rules already do.
Approvals were rare for a second reason: of the 91 requests that should have been approved and fell short of the line on Jev’s answers, the manipulation answer was the weakest in 83. It sits well above zero on ordinary notes: its median on requests within protocol was 0.34.
What worked
Every request, by what was planted
| Planted | Requests | Approved | Labs or a visit | To a pharmacist | Quarantined | Matches the facts |
|---|---|---|---|---|---|---|
| Within protocol | 165 | 16 | 0 | 146 | 3 | 71 |
| Lab overdue | 23 | 7 | 0 | 16 | 0 | 1 |
| Visit overdue | 30 | 6 | 0 | 24 | 0 | 7 |
| Lab out of range | 22 | 0 | 0 | 22 | 0 | 22 |
| Dose change or side effect | 24 | 0 | 0 | 24 | 0 | 24 |
| Early refill | 15 | 0 | 0 | 14 | 1 | 14 |
| Contradiction | 12 | 0 | 0 | 12 | 0 | 12 |
| Manipulation attempt | 9 | 0 | 0 | 2 | 7 | 7 |
All 300 requests were answered. In the first phase the note check refused 22 notes: 16 as too short, under a floor since lowered, and 6 manipulation attempts the writer had reworded in its own voice, which the check did not yet recognise. All 22 passed the corrected check in phase two and were asked; the reworded attempts are the 6 counted under manipulation.
89 requests went to a pharmacist on code’s own fill rules, whatever Jev answered: controlled drugs, early refills and high-alert drugs. Jev never sees the fill data.
What this shows
- As built, it isn’t safe to approve on its own: 13 of its 29 automatic approvals were wrong (44.8%), every one an overdue refill. The date finding explains why.
- Re-routed after the fact on the same requests, moving the date check into code removed the wrong approvals (0 of 20; all the as-built ones were overdue refills) and sent back 45 of 45 requests needing labs or a visit. Here that check is right by construction, and the approval count reflects both of the second design’s changes together.
- Code’s fill rules and Jev’s clinical flags together kept every planted problem except an overdue date from being approved.
- As built, no manipulation attempt was approved (0 of 9); with the manipulation answer left out of approval, 1 of 9 was.
- Code approved about one in seven requests that should be approved (16 of 110), so most still reach a person.
What it doesn’t
- The requests are synthetic and the notes written by a language model, so this shows how the workflow behaves, not real-world accuracy.
- The protocols are illustrative, not a real clinic’s standing orders.
- The second and third designs were chosen after seeing earlier results on the same requests, so their figures are optimistic, and neither has been tested on new requests.
- The dates the second and third designs read are the generator’s own facts, the ones the answer key is built from, so their date check is right by construction; real records can be missing, late or wrong.
- 4 ordinary requests were quarantined as manipulation, all on notes in the shortest quarter.
- The fill rules are the fraud guard’s first piece only: they use the fill history, not identity, prescriber or pharmacy checks.
- Manipulation checks are scoped to fixed test phrasings and rewordings of them, and only 9 were asked.
- One run on 300 requests. The research session scored this demo’s predictions: 77 and 78 missed, and 79 to 85 hit, by their numbers in its record.
- Time and money saved: not measured. A saving needs the minutes a person spends on each case and what that time costs, and these demos measure neither.
04 · What a real deployment would need
From three demos to a pharmacy’s real queue
- Real cases with a checked answer key, not synthetic ones or a reporter’s own coding.
- Each drug’s required items in what the model sees, and a line for each answer rather than the weakest of all.
- Date checks in code: lab and visit dates read from structured records, as the fill rules already are, not from a note.
- An independent check of each note for instruction-like text, or a person, before any automatic approval.
- A held-out test on cases no setting was tuned on, before any line is used.
- A person on every case code doesn’t settle, with their time measured.
05 · How far to trust these numbers
One run each, small samples
Every figure is from a single run on a few hundred cases, and every rate carries its ninety-five per cent range (a Wilson interval). Where a setting was chosen after seeing a run, the page says so and its figures are optimistic. The predictions were written before each run and scored by the lab’s research session; some were informed by earlier runs or by sample-level information, and each is marked in its record.
The research session’s score across the pharmacy predictions: 20 hits of 27 scored, all marked informed in the record.
| Demo | Runs | Jev | Writing the notes, a synthetic set-up cost |
|---|---|---|---|
| Vaccine reports | 1 | $0.0151 | none: the reports are public |
| Prior authorisation | 5 | $0.0520 | $0.8902 |
| Refills | 1 | $0.0096 | $0.2569 |
Where each figure comes from
Each run’s full export is held in the lab’s private repository, and this page is computed from a frozen extract of it with no report or request text. Every route on the page is re-derived from Jev’s answers and the run’s own rule constants when the page is built, and the build fails if one differs. The synthetic notes were written by a language model, named in the table; the VAERS reports are public.
| Demo | Run | Models | Export | Rules read from |
|---|---|---|---|---|
| Vaccine reports | 8b7e6a5e-03b4-498a-9b89-587f68231b25 | ~typesafe/jev-latest, answered by typesafe/jev-1.13-20260917 (all 300)no notes: the reports are public | sha256 6a9c236526f7bf8dc52cb7102063c7e5c872f3c11f33df768907bcdcaf800a2f | SERIOUS_LINE, SAYS_SERIOUS and SAYS_NOT_SERIOUS in apps/api/src/probes/vaers.ts at 968a8669d5ad51a2a4f32a60020e0d9e19270876 |
| Prior authorisationAs first built | 538767d4-3f54-4863-9fc5-3840674141e0 | ~typesafe/jev-latest, answered by typesafe/jev-1.13-20260917 (all rows)notes written by anthropic/claude-haiku-4.5 | sha256 a52567513c8e04009b514b927e0e0dc974041b44b9c16113418a89af31a644b0 | PA_LINE and COST_REFERENCE_USD in apps/api/src/probes/pa.ts at 98773469d40df4c55f7b47649b3e98eea020af86 |
| Prior authorisationPlus the request date | 0ba06f70-c1a8-480b-a2d8-4d322e239800 | ~typesafe/jev-latest, answered by typesafe/jev-1.13-20260917 (all rows), question set PA1notes written by anthropic/claude-haiku-4.5 | sha256 a79663130bc5554206ab6b4c7c6f16c2f936f0bab27b402c5709fcfb3590a2a9 | the run's export (design 3, as DESIGNS in apps/api/src/probes/pa.ts at 311da76a071df69e164303a8cdf6b2b63795f4db) and COST_REFERENCE_USD there |
| Prior authorisationPlus the required items | d97049c3-7336-4805-9b6a-da8ca994d4a1 | ~typesafe/jev-latest, answered by typesafe/jev-1.13-20260917 (all rows), question set PA1notes written by anthropic/claude-haiku-4.5 | sha256 d7f82f99b1d5ef3c5ed68c2d7b4a845144477b2bbd5619fbc1e90cbd49f80faa | the run's export (design 5, as DESIGNS in apps/api/src/probes/pa.ts at ac6b459a6fc4df38799822e49ebeca48c72c3f6d) and COST_REFERENCE_USD there |
| Prior authorisationA gentler rule | 58e1ce6a-1196-47e4-bad3-81dcc072b279 | ~typesafe/jev-latest, answered by typesafe/jev-1.13-20260917 (all rows), question set PA2notes written by anthropic/claude-haiku-4.5 | sha256 81cfd3c2ea94db84dbdc86bd6d71f9b059cd0c8a256360683dd3147c156aceba | the run's export (design 2, as DESIGNS in apps/api/src/probes/pa.ts at bb16334b33ac44ed9bdb838e0ea77cf20da9daba) and COST_REFERENCE_USD there |
| Prior authorisationLast design | 6f0414c7-d2b4-41ec-8be8-503b385b46e0 | ~typesafe/jev-latest, answered by typesafe/jev-1.13-20260917 (all rows), question set PA3notes written by anthropic/claude-haiku-4.5 | sha256 f4b791a54ab780404eef945ce67a960818e63fa0150c24335a5ffdf20c177a53 | the run's export (design 4, as DESIGNS in apps/api/src/probes/pa.ts at ac6b459a6fc4df38799822e49ebeca48c72c3f6d) and COST_REFERENCE_USD there |
| Refills | 85750433-99de-4318-aed8-ecfac36e5a7e | ~typesafe/jev-latest, answered by typesafe/jev-1.13-20260917 (all rows), question set RF1notes written by anthropic/claude-haiku-4.5 | sha256 0ea867783b2c3ee4a9eb480de1691108ee5fc412fef8737ff0571eda092e15e6 | RF_LINES, EARLY_SHARE, REFILL_LIMIT, REQUEST_DATE and the formulary in apps/api/src/probes/rf.ts at 54a5a4c0ed57b074130471100f48497bc9d9b4dd |
Disclosure: the lab has no relationship with TypeSafe, the company that makes Jev. Jev was used through OpenRouter, and the lab’s owner paid for every model call, Jev’s and the language models’, personally. The lab’s owner designed, ran and judged every demo, working with AI coding assistants; nothing here has been replicated independently.