CorrectMED AI
How CorrectMED AI is measured
Reading is measured on documents whose answers are written down before any reading, and every figure comes with its sample size.
The reader
Everything CorrectMED reads goes through the CorrectMED AI reader cm-reader-1. It copies printed values, marks each one clear, unclear or absent, and boxes where it came from. Fixed rules, not AI, compare the values, and people decide.
The exact model version is given to reviewers on request.
How it's measured
- Generated delivery notes. Our generator draws delivery notes with fictional suppliers and generic products, and writes down every printed value and where it sits on the page before CorrectMED AI sees the image. Some fields are made unreadable on purpose: cut off, faded or smudged. The scored notes use layouts and fonts that were never used while tuning, and 17 of the 30 get a phone photo look (a slight turn, blur and shading). They're generated fictional notes, not photos of real deliveries.
- Alert matching rules. Rule tests, not accuracy: lines built from real NAFDAC alert rows, each with the outcome the rules must give, and random lines with no connection to any alert are checked against the alert index, counting matches the rules should never make. No AI is involved.
What's counted
- Readable fields copied exactly, character for character.
- Unreadable fields wrongly marked clear. This is the number that matters most: a value marked clear reaches the rules.
- Lines missed, merged into another line, or read when they aren't on the page.
- Source boxes on the right text, and how much they overlap the true box.
- Time per read, as the median and the 95th percentile.
Every figure shows its count and a Wilson 95% interval, so you can see how much a small sample can tell you.
Results
Generated
Generated delivery notes
30 generated fictional delivery notes, not photos of real deliveries, in layouts and fonts never used while tuning, each read once through the app's own read path with the settings fixed first. 120 fields were made unreadable on purpose (40 cut off, 40 faded, 40 smudged), 30 were faded but still legible and 27 cells were left blank. Source boxes overlapped the true box by 0.83 on average (IoU). Sample: 30 notes, 300 lines, 2,400 fields.
- Readable fields copied exactly
- 2,220 of 2,223 (99.9%)95% interval 99.6% to 100.0%
- Unreadable fields wrongly marked clear
- 0 of 120 (0%)95% interval 0% to 3.1%
- Lines missed
- 0 of 300 (0%)95% interval 0% to 1.3%
- Lines merged into another line
- 0 of 300 (0%)95% interval 0% to 1.3%
- Lines read that aren't on the page
- 0 of 300 (0%)95% interval 0% to 1.3%
- Source boxes centred on the right text
- 2,253 of 2,253 (100%)95% interval 99.8% to 100.0%
- Values the rules parse the same as the printed value
- 2,221 of 2,222 (100.0%)95% interval 99.7% to 100.0%
- Blank cells marked absent
- 27 of 27 (100%)95% interval 87.5% to 100%
- Faded but legible fields copied and marked clear
- 16 of 30 (53.3%)95% interval 36.1% to 69.8%
- Supplier, note number and date copied exactly
- 84 of 90 (93.3%)95% interval 86.2% to 96.9%
- Seconds per read
- median 16.0, 95th percentile 20.5
Failures in this run
- One quantity was read as 40 where the note printed 41, and marked clear. It's the one value in the run that reaches the rules wrong; the staff count is compared against it.
- One quantity, 4, was read correctly but marked unclear, so it shows as Could not read.
- One pack size printed as "1 X 6" was copied as "1X6". The rules read the same pack size, but it isn't an exact copy.
- 14 of the 30 faded but legible fields were marked unclear: 13 with the right value beside them and 1 left empty. None was read wrong.
- On 6 notes the note number was copied without its "No." prefix.
- The backup reader started on 26 of 30 reads, because the first answer took longer than 12 seconds. The first reader still answered first every time, so no read failed and none ran past 21 seconds.
Generated
Alert matching rules (rule tests, not accuracy)
260 lines checked against an index of 360 NAFDAC alert rows: 60 built from real rows, each with the outcome the rules must give (an exact copy, a look alike or one character batch change, another product, another strength, no batch, an unrelated batch), and 200 random lines from the sample generator with no connection to any alert. For this test every row counts as reviewed, so the exact match rules are exercised; in the app a row gives an exact match only once a person has reviewed it. No AI is involved. Sample: 260 lines, 360 alert rows.
- Exact matches the rules should never make
- 0 of 260 (0%)95% interval 0% to 1.5%
- Built lines given the outcome they were built for
- 60 of 60 (100%)95% interval 94.0% to 100.0%
- Exact copies of an alert row given an exact match
- 9 of 9 (100%)95% interval 70.1% to 100%
- Random lines given an uncertain match
- 0 of 200 (0%)95% interval 0% to 1.9%
Failures in this run
- No failures. The rules keep dosage forms apart, so an artemether/lumefantrine tablet line doesn't meet the two NAFDAC notices that cover every batch of the dry powder suspension.
Limits
- The scored notes are generated fictional notes, not photos of real deliveries. They show how reading behaves on printed notes in known layouts, not on a month of real deliveries, handwriting or crumpled paper.
- Each scored set is read once, with prompts and settings fixed first. A new comparison uses a newly generated set.
- Labels for generated notes are written by the generator before any reading.
- One run is one sample: answers can vary between runs even with the same settings.
- Timings come from the machine that ran the evaluation, one read at a time. The live app runs elsewhere, so its times differ.
What this doesn't measure
There's no detection rate here. CorrectMED reads documents and packs, applies fixed rules and shows published NAFDAC alerts. It doesn't judge any medicine, so there's no rate of that kind to report.
You can try the reader yourself on a fresh note from the samples page, and read how CorrectMED AI is used.