An AI scribe is usually sold on time saved, and time saved is measurable. This guide sets out what to capture before and after a trial, how to run a fair two-week baseline, and which numbers look like evidence while telling you nothing. Nearly all of it applies to any ambient scribe, whichever vendor a practice is considering.
Decide what a result looks like before you start
Most trials of clinical software end with an impression. A clinician uses the tool for a fortnight, reports that the notes feel faster, and the practice decides on that feeling. Impressions matter, because a tool clinicians dislike will not survive contact with a busy list, but a feeling is a weak basis for a decision that touches every consult in the practice.
The stronger approach costs almost nothing. Agree in advance what a successful trial would show, capture a baseline over two normal weeks, run the trial under the same conditions, and compare like with like. The question has two parts: how much clinician time the scribe returns, and whether the record is at least as good as it was before. A trial that answers only the first is incomplete.
It also helps to decide up front where any saved time should go. Some practices want clinicians finishing on time. Some want capacity for an extra appointment each session. Some simply want letters out faster. These are different outcomes, and the measurement that proves one will not automatically prove another.
The four measures that matter
Four measures cover most of the real value of an ambient scribe, and each can be captured from systems the practice already runs. The first two measure clinician time directly. Minutes per note must include review, because a draft is not a clinical record until a clinician has read it, corrected it and signed it. A scribe that produces a draft in seconds but needs ten minutes of correction has saved far less than it appears to.
The second two measure the health of the record. Letter turnaround is a practice-level number that patients and referrers feel directly. Recall and follow-up accuracy is the easiest to overlook and arguably the most important. A note written late in the evening from memory tends to drop actions that a note drafted from the consult itself tends to keep. Audit a sample of signed notes from both periods and compare what actually made it into the recall system.
- Minutes per note. The time from the end of the consult to the note being reviewed and signed in the practice software.
- After-hours documentation time. Notes, letters and results follow-up completed outside booked sessions, including evenings and weekends.
- Letter turnaround. The days between the consult and the referral or specialist letter leaving the practice.
- Recall and follow-up accuracy. Whether the actions agreed in the consult (recalls, referrals, medication changes, results to chase) landed correctly in the record and the recall system.
How to run a fair two-week baseline
Pick two ordinary weeks. Avoid school holidays, public holidays, conference leave and any week where a clinician is covering someone else's list. The baseline needs to look like the practice's normal life, because that is what the trial period will be compared against.
Measure the clinicians who will actually trial the scribe, ideally two or three with different documentation styles: one fast typist, one who writes notes late, one somewhere in between. A baseline built only on the practice's most efficient documenter will make any tool look unnecessary, and one built only on the most burdened will make any tool look miraculous.
Keep the collection light. Practice software already timestamps when notes are signed, so the gap between the last appointment and the last signed note each day is retrievable without anyone holding a stopwatch. Add a short end-of-day tally from each participating clinician: how many notes were left unsigned at close, and roughly how long documentation took after the last patient. Pull letter dates from the correspondence log. None of it needs more than a spreadsheet and the reports the software already produces.
Tell the team what is being measured and why. The measurement is a test of the tool. If clinicians suspect it is quietly a test of them, the numbers will be gamed and the trial wasted.
Run the trial on the same terms
The trial period should mirror the baseline: the same clinicians, a similar case mix, two ordinary weeks. Every new tool has a learning curve, and the first few days will be slower while clinicians settle the consent script and build the review habit. Either run the trial long enough for the early days to wash out, or record both weeks and weight the second more heavily when you compare.
Measure the same things the same way. If the baseline counted minutes from consult end to signed note, the trial must too, including every minute the clinician spends reading and correcting the draft. If the baseline counted unsigned notes at close of business, count them again. Resist the temptation to add flattering new measures mid-trial.
Add one quality check the baseline could not have: a correction log. Ask each clinician to flag drafts that needed substantial rework and what kind (a wrong medication name, a missed negative finding, words attributed to the wrong speaker). A scribe with a modest time saving and a low correction rate can be a better buy than one with a large time saving and a correction pattern that erodes trust in the record.
Vanity metrics to ignore
Some numbers look like evidence and measure nothing a practice should pay for. The most common is volume: notes generated, minutes of audio transcribed, words produced. Volume measures usage. It says nothing about whether a clinician got home earlier or whether the record improved.
Time-to-draft is the second trap. A draft appearing seconds after the consult is genuinely useful, but quoting it as the documentation time ignores the review that turns a draft into a record. The only end-to-end number worth comparing is consult end to signed note.
Hours-saved self-estimates collected without a baseline are the third. People estimate their own time use poorly, and a clinician enjoying a new tool will overestimate the saving in complete good faith. That is exactly why the two-week baseline exists: it replaces recollection with timestamps.
Satisfaction still matters. A tool the team likes gets used, and a tool they resent gets abandoned regardless of what it saves. Collect satisfaction alongside the timing data, and treat it as a condition of success rather than a substitute for the numbers.
How aurii fits a measured trial
This section is about our product. Everything above is not.
In a trial, aurii is straightforward to measure because the workflow separates cleanly into the parts described above. The draft consult note is ready when the consult ends, so the clinician's remaining time cost is review and sign-off, and the minutes-per-note measure reads directly off the practice software timestamps. Letters and discharge summaries draft from the same consult, which pulls letter turnaround into the same day's work rather than a Friday backlog.
Nothing becomes part of the record until a clinician has reviewed and signed it, which keeps the accuracy measures meaningful: what you audit in the trial period is clinician-approved output, held to the same standard as the baseline notes. Consults are captured with the patient's consent, and everything is transcribed and stored in Australia.
Run the baseline exactly as described before switching anything on, then trial with the same clinicians for two weeks and compare. The measures in this guide apply to any ambient scribe, and a fair trial run against a real baseline will show what any of them, including aurii, is worth to your practice.
Common questions
Two ordinary weeks is enough for most practices. Avoid holiday periods and unusual rosters, and measure the same clinicians who will run the trial. The goal is a picture of normal documentation load, captured from timestamps the practice software already records.
Yes, always. A draft is not a clinical record until a clinician has read, corrected and signed it, so the honest measure runs from the end of the consult to the signed note. Quoting draft-generation time alone flatters any scribe.
Any number that grows with usage rather than with benefit: notes generated, minutes transcribed, words produced. Hours-saved estimates collected without a baseline belong in the same category, because self-reported time savings are unreliable however sincere they are.
Use timestamps the systems already keep. The gap between the last appointment and the last signed note, plus a count of notes left unsigned at close of business, captures the after-hours pattern without anyone watching over shoulders. A brief end-of-day self-report fills the gaps.
No. Two or three clinicians with different documentation styles give a fairer picture than one enthusiast, and a whole-of-practice rollout before the numbers are in adds risk without adding information. Expand once the trial supports the decision.
Check where the time went before calling it small. A modest per-note saving can still remove most after-hours documentation or shorten letter turnaround considerably, and those may be the outcomes the practice actually wanted. This is why agreeing the target outcome before the trial matters.
This is general information about evaluating documentation tools. It is not clinical, legal or financial advice. More guides sit on the resources hub. If your practice needs a question answered before it adopts AI documentation, tell us and we will write it: hello@aurii.com.au.