Freight brokerage demo
Can a load move from tender to invoice with the routine work off a broker’s desk?
A freight broker’s team takes a shipper’s tender, books a carrier, tracks the truck, checks the proof of delivery and pays the invoice. Much of that is routine. This demo follows synthetic loads through every step and adds automation one rung at a time: fixed rules, then Jev, a narrow decision model from TypeSafe that answers typed questions, then a wide model that reads and drafts documents. People keep every exception, and nothing is refused automatically.
Demo, not validated · synthetic loads and messages
Results pending: development runs under way; a test on new loads, declared in advance, comes after.
The demo, as it stands
- Tender
- Carrier reply
- Status
- Proof of delivery
- Invoice
01 · One load, step by step
What arrives at each step, and what should happen?
Demo, not validated · synthetic loads and messages
Each load has five steps. The status step has two checkpoints, a ping from the truck and a message from the carrier, so a load has six step-instances. Each load’s facts, and what should happen at every step, are its answer key: stored with the run and never shown to a model.
| Step | What arrives | What should happen |
|---|---|---|
| Tender | The shipper’s tender: a structured message or an email. | A correct load record, or a person when the tender is incomplete, unclear or a change to an existing load. |
| Carrier reply | The carrier’s answer to the rate offer, by email. | Book a clean yes at the offered rate; a person for a counter-offer, an added charge, a change of date or equipment, or a change of payment or carrier.How stale is an old rate? The earlier study of published rates is now a footnote to this step: the rate-staleness study. |
| Status: the truck | A location ping from the truck’s logging device. | An automatic update to the shipper, or a person when the truck is late or has stopped. |
| Status: the carrier | The carrier’s or driver’s message. | An automatic update, or the shipper notified, a tracker, a security hold or quarantine. |
| Proof of delivery | An electronic proof of delivery, or a scanned one’s text. | Accept a clean one; a person for shortage or damage, a missing signature, a date that doesn’t reconcile, or a refusal. |
| Invoice | A structured invoice or an emailed one. | Approve payment when it matches the rate confirmation plus approved charges; a person otherwise. |
02 · The ladder
Who does each step, at each rung?
The same loads run through four rungs. Each rung adds one kind of automation and keeps everything the rung below it does. Code decides at every rung: it applies the rules, compares every amount and date, and sends anything unclear to a person.
- L0 · All people. The reference: a person does every step.
- L1 · Fixed rules. Code handles what is structured; free text goes to a person.
- L2 · Plus Jev. Jev answers typed questions about replies, status messages and proof texts.
- L3 · Plus the wide model. The wide model extracts fields from emails and drafts documents; code checks the dates, lane, equipment and amounts it gives, but does not check party names or invoice numbers against the email.
Key
- Code (square)
- Jev, a narrow decision model (circle)
- The wide model, which reads and drafts (diamond)
- A person (triangle)
A routine load at the top rung
Each stop is drawn in the shape of whoever acts there, as in the ladder. Code decides every route; anything unclear goes to a person.
The ladder as a table
| Step | L0 · All people | L1 · Fixed rules | L2 · Plus Jev | L3 · Plus the wide model |
|---|---|---|---|---|
| Tender | A person keys it. | Code reads a structured tender; an email goes to a person. | As fixed rules. Jev checks an emailed tender only for an instruction to the system. | The wide model extracts an emailed tender into the record; code checks its dates, lane and equipment, and any failure goes to a person. Jev checks for an instruction to the system. |
| Carrier reply | A person reads it. | To a person: a simple check for a yes never books. | Jev answers typed questions about the reply; code compares every amount with the offer and books a clean yes. | As with Jev, and the wide model drafts the rate confirmation; code inserts every number and date. |
| Status: the truck | A person calls. | Code’s lateness and stop rules; a late or stopped truck goes to a person. | As fixed rules. | As fixed rules. |
| Status: the carrier | A person calls. | To a person: code can’t read the message. | Jev answers typed questions about the message, and code sends it on. | As with Jev, and the wide model drafts the notice to the shipper; code inserts every time and figure. |
| Proof of delivery | A person checks it. | Code reconciles an electronic proof with the booking; a text goes to a person. | Jev answers typed questions about a proof’s text; with code’s date check, a clean one is accepted. | As with Jev, and the wide model extracts the delivered date, time and pieces for code to reconcile. |
| Invoice | A person matches it. | Code matches a structured invoice to the rate confirmation; an email goes to a person. | As fixed rules. Jev checks an emailed invoice only for an instruction to the system. | The wide model extracts the line items and code sums and matches them; payment needs the amounts to match exactly and a clean proof of delivery, otherwise a person. Jev checks for an instruction to the system. |
- The wide model is used only where it reads or writes what Jev can’t: extracting fields and drafting. It is never a second opinion on Jev’s reading of the same text.
- Dates, times and money are handled in code. A figure the wide model extracts is never trusted on its own: code checks it against the load’s other records, and any mismatch goes to a person.
- A cheap check can send work to a person, but never books, approves or pays.
- An answer that a message carries an instruction to the system only ever quarantines it.
03 · What is measured
What will each rung take off the desk, and what will it get wrong?
- Closed with no person: step-instances completed with no person.
- Wrong automation: an automated step-instance whose action differs from the answer key, counted over the automated ones.
- Missed exceptions: step-instances the key sends to a person that were automated instead: the costly error.
- Per load: loads no person touched end to end, and person touches per load.
- Cost and speed: paid model cost per load, apart from writing the synthetic inputs, and the automated path’s time per load.
All people is the reference. Its errors are not measured: nobody’s error rate is assumed, and the page never shows them as zero.
| Rung | Closed with no person | Wrong automation | Missed exceptions | Loads with no person | Touches per load | Model cost per load |
|---|---|---|---|---|---|---|
| L0 · All people | Pending | Not measured | Not measured | Pending | Pending | Pending |
| L1 · Fixed rules | Pending | Pending | Pending | Pending | Pending | Pending |
| L2 · Plus Jev | Pending | Pending | Pending | Pending | Pending | Pending |
| L3 · Plus the wide model | Pending | Pending | Pending | Pending | Pending | Pending |
Every figure here comes from the run’s extract, and appears only once the lab’s research session has checked how it was measured. Until then each cell says pending.
04 · The status step, run on its own
How did the status step do on its own?
Demo, not validated · synthetic loads and messages
This is the status step run on its own, before the ladder, at its line of 0.5 on every answer: not a ladder result, and not the line the flow will fit.
The carrier’s and driver’s messages were read on the development loads by an earlier demo. Jev answered typed questions about each message, and code sent each load on. Code alone sees only the truck’s ping and the timetable.
Instructions to the system: 9 of 9 were quarantined. Jev flagged 8 and the separate screen 7, and between them they caught 9. On the instructions to the system: they reword pharmacy phrasings and are out of domain for logistics, so the result is likely optimistic.
Loads with no planted instruction that were quarantined: 2, the security requests noted above. Loads excluded under the fixed exclusion rule: 0.
Jev’s cost: about $0.00003 a load ($0.0079 in all); writing the synthetic messages, a set-up cost, $0.1879.
05 · What was predicted
The ladder: registered before any flow run, not yet scored
| Prediction | As registered | Confidence | Result |
|---|---|---|---|
| 99 | The share of step-instances closed with no person rises at every rung: L1, then L2, then L3. | 85% | Not yet scored |
| 100 | L3's gain over L2, in step-instances closed with no person, is larger than L2's gain over L1. | 60% | Not yet scored |
| 101 | At L3, the Wilson 95% upper bound on wrong automation, pooled over all automated step-instances, stays at or below 2%, and no planted security flag at any step (reply, invoice or status message) is automated. | 45% | Not yet scored |
| 102 | At L2 and at L3, no more than 5% of exception step-instances are automated. | 55% | Not yet scored |
| 103 | At L3, the paid model cost is under $0.02 a load, and the median automated path under 20 seconds a load. | 65% | Not yet scored |
- Registered on 30 September 2026, before any flow run. The texts were fixed before the status step's development results arrived, and none was changed after.
- Informed: before the models were asked anything, the research session read the writing phase's figures, a scan of every development text, and a sample beside its facts.
- Two design changes were made after registration, before the models were asked anything, and both make prediction 101 easier to meet: a narrow-model check for instructions to the system on emailed tenders and invoices, and a fixed line for every manipulation and security answer instead of a fitted one.
- Prediction 103's speed is measured over every load, counting every model call made for it, whether or not its steps closed.
- Not yet scored. They are scored once, on 600 held-out loads, after the development lines are frozen.
The status step: scored
| Prediction | As registered | Confidence | Measured | Result |
|---|---|---|---|---|
| 95 | At most 5% of the loads that the facts send to a person are updated automatically. | 50% | 0.0% | Hit |
| 96 | At least 70% of the loads that need no person are updated automatically. | 50% | 100.0% | HitAs disclosed before any message existed (#236), these messages were kept free of every other scenario's words, so this is likely optimistic. |
| 97 | Every planted security red flag but at most one reaches a person. | 60% | 100.0% | HitTwo of the fifteen, both requests to pay another company, went to quarantine rather than the security desk; both reached a person. |
| 98 | Of the loads code alone would update automatically but the facts send to a person, Jev's answers send at least 80% to a person. | 55% | 100.0% | Hit |
Informed: the research session read phase 1's figures, a scan of all 300 messages and a 20-message sample before Jev was asked anything.
The scores are the lab’s research session’s, not this page’s. For scale: across all the lab’s demos, the research session’s score stands at 57 hits of 80 predictions scored.
06 · Where the demo stands
Where does the demo stand?
| Phase | What it does | State |
|---|---|---|
| F1 | The generator, each load’s answer key and the fixed rules. | Done |
| F2 | Writing the tenders, replies, proof texts and invoices, then code’s check of every text, and the research session’s read of a sample. | Done |
| F3 | Development: Jev’s and the wide model’s reads, every route, and the lines fitted without any load seeing its own fit. | Run; results under review |
| Freeze | The lines, the prompts and the models frozen together, before the test is drawn. | Later |
| F4 | The test: drawn, written, checked and run once. | Later |
The development set is 300 loads; the test runs once on 600 new loads.
07 · How far to trust it
Synthetic loads, declared rules
Every load is synthetic, on real lanes, and every message is written by a language model from the load’s facts, then checked by code. The answer key is the facts, not a person’s judgement. The lines are fitted on the development loads, then frozen and tested once on new loads declared before they were drawn. Nothing here is validated.
How it was built
- The structured documents (the structured tenders and invoices, and the electronic proofs of delivery) use the lab’s own formats, so they measure routing, not parsing.
- Every booked carrier is treated as eligible: vetting carriers is out of scope.
- The instructions to the system were written by the builder, who knew the screen’s patterns.
- The development status messages use earlier instructions worded for pharmacy, out of domain for freight.
- The proof-of-delivery date check reads the lab’s own date format.
Time and money saved: not measured. A saving needs the minutes a person spends on each step and what that time costs, and this demo measures neither.
Disclosure: the lab has no relationship with TypeSafe, the company that makes Jev. Jev was used through OpenRouter, and the lab’s owner paid for every model call, Jev’s and the language models’, personally. The lab’s owner designed, ran and judged this demo, working with AI coding assistants; nothing here has been replicated independently.