Vertex had a working demo and no way to prove it was safe. We built the evaluation harness first, then the product — and the acceptance rate went from 61% to 96%.
Up from 61% on the original prototype.
Median minutes per consultation note, measured over 12 weeks.
On the weighted clinical scorer, with zero critical medication errors.
Passed independent audit at first attempt, including the AI-specific controls.
A prototype note-drafting assistant impressed in demos and failed in clinics. Clinicians rejected roughly two in five drafts, usually for subtly wrong medication detail. Nobody could say whether a prompt change made things better or worse, because there was no way to measure it.
We built a 2,400-case gold-standard dataset with clinician-authored reference notes, and a scorer weighted so a medication error counted forty times a stylistic one.
Every clinical claim in a draft had to cite a span from the transcript or the patient record. Ungrounded sentences were suppressed rather than shown, which cut hallucinated detail to near zero.
Drafts arrive with citations inline and low-confidence spans highlighted. Clinicians correct rather than accept-or-reject, and every correction feeds the eval set.
No prompt, model or retrieval change ships without beating the current scorecard. Six candidate changes were rejected on that basis during the build.
“They refused to build features for two months while they built the measuring stick. It was frustrating at the time and it is the only reason this thing is in clinics today.”
The first conversation is with the engineers who would run your engagement, not an account manager.
hello@cymbiote.com