Ambient scribes handle the most sensitive material a practice holds, and the training question is the one suppliers answer least plainly. This article separates using a consultation to produce your note from using it to improve a model, works through why a de-identification claim rarely settles it, sets out the secondary use rule and the model provider sitting behind the vendor, covers what a tuned model does to that analysis, and gives the wording to require in the agreement. Most of it applies to any ambient tool. One section covers how aurii answers it.
The difference between processing a consultation and training on it
Every ambient scribe processes the consultation. Audio moves through speech recognition, a language model drafts against a note template, the draft comes back to the clinician, and copies of the audio, the transcript and the draft persist somewhere for some period. Training is a different activity with a different beneficiary. It means the content of your consultations is retained as material used to change how a model behaves, so the product performs differently in future for every practice that uses it.
A wide band of activity sits between the two, and vendor answers routinely blur it. Product improvement, quality assurance, evaluation, human review of samples, prompt and template tuning against real transcripts, fine-tuning a model for a single customer, and aggregate analytics all involve keeping and reading consultation content, and none of them has to be described as training. A supplier can say truthfully that it does not train its models on customer data while a review team reads transcripts every week and a set of real Australian consultations sits in a test harness.
The question that produces a usable answer is therefore narrower than whether the vendor trains on your data. Ask what copies of the audio, transcript and note exist once the clinician signs, which people and which automated processes can read each copy, for what purposes, and for how long. Framed that way the supplier has to enumerate, and the list is what you can hold them to later. These are the uses worth naming one at a time:
- Producing this practice's notes, letters and summaries.
- Fault diagnosis by support staff, including whether that access is logged and whether the practice is told.
- Human review of samples for quality, including who reviews and whether the practice can decline it.
- Evaluation, benchmarking and regression testing against stored real consultations.
- Training, fine-tuning or adapting any model, the supplier's own or a third party's.
- Anything described as aggregated, statistical or de-identified, stated separately with the method used.
Why a de-identification claim does not settle the question
The Privacy Act 1988 applies to personal information, which it defines as information or an opinion about an identified individual, or an individual who is reasonably identifiable. A de-identification claim is a claim that the second limb no longer holds. Stripping the name, the date of birth and the Medicare number is the start of that work and nowhere near the end, because singling a person out needs no formal identifier, only enough detail that someone holding other information can do it.
De-identification describes data in a particular context, and it is not a permanent property of a file. The same transcript can be de-identified in the hands of a party holding nothing else, and personal information in the hands of a party holding an appointment book, a billing extract or a pathology feed. The OAIC's guidance on developing and training generative AI models puts it plainly: de-identification is context dependent, may be difficult to achieve, and developers should treat data as personal information wherever there is doubt. A claim about de-identified training data therefore has to name the release context it was assessed against and who assessed it.
The exposure sits with the practice. The practice collected the health information, the practice chose the tool, and if the de-identification turns out to be weaker than claimed, an unauthorised disclosure of health information likely to result in serious harm is an eligible data breach under the Notifiable Data Breaches scheme. The first question put to the practice will be what it knew about the arrangement when it signed.
What survives redaction in consultation audio and clinical free text
Speech carries identity independently of what is said. A recorded voice is a biometric characteristic, and two recordings of the same person can be matched by software that never processes a word of the content, so removing names from the transcript does nothing to the speaker in the waveform. Unless a supplier discards the audio and retains text only, a description of retained consultation audio as de-identified will not survive examination, so ask whether the audio is discarded and at what point.
Clinical free text carries the same problem in a different form. Patients name their partner, their employer, their street, the pharmacy they use and the school their child attends. Clinicians dictate the referring doctor's name and the hospital where the operation happened, and dates of injury, surgery and admission work as keys into other datasets. A rare diagnosis, an age and a regional town can be unique in Australia with no formal identifier present anywhere in the record. Automated redaction is itself a statistical model with an error rate, running across every consultation a vendor processes, so residual identifiers accumulate in proportion to volume.
Measurement is therefore the thing to ask about. Ask how the redaction was tested, against what set of real consultations, and whether Australian personal names, place names and drug brand names were in that set. A vendor that has done the work answers with a method and a recall figure, and one that has not answers with the categories alone.
Language models can also reproduce distinctive strings from their training data, which concentrates the exposure in the unusual record, and clinical records hold a great many unusual records. A consultation captures people who were never asked as well: the family member who speaks, the interpreter, the student sitting in, the clinician's own voice, and in a shared space the conversation next door. Consent to record so the doctor can write an accurate note does not stretch to any of them becoming training material.
The secondary use rule, and where vendor training sits outside it
Health information is sensitive information under the Privacy Act, and the small business exemption is no help here, because an organisation that provides a health service and holds health information is covered whatever its turnover. Practices in New South Wales, Victoria and the Australian Capital Territory also sit under state and territory health records legislation. A practice collects health information for a primary purpose, providing care and keeping the clinical record, and using it for anything else is a secondary use governed by Australian Privacy Principle 6. That leaves consent, or a use the patient would reasonably expect that is directly related to the primary purpose, since the direct test is the one that applies to sensitive information, plus a short list of statutory exceptions.
Improving a supplier's commercial model is not directly related to treating the patient in the room, and no patient in an ordinary conversation about note-taking expects it. Australian privacy law does contain a route for handling health information without consent for research relevant to public health or public safety, through the permitted health situations in section 16B, which require the research to follow guidelines approved under section 95A of the Act and to go before a human research ethics committee. That committee weighs the public interest in the research against the public interest in privacy, and supplier product development does not reach it.
The part that catches practices is who owes the obligation. The practice is the entity that collected the information, and the supplier is normally handling it on the practice's behalf. When a supplier takes that content for its own purposes, the use is happening because the practice accepted a term in a supply agreement. A term in a supply agreement is not consent from the patient, who signed nothing.
The OAIC's generative AI guidance reaches the same place from the developer's side: where a developer cannot clearly establish that a secondary use for an AI-related purpose was within reasonable expectations and related to a primary purpose, it should seek consent for that use, offer a meaningful and informed ability to opt out, or both. If a practice does permit the use, the disclosure obligation lands on the practice too. Australian Privacy Principle 5 requires patients to be told at collection what their information will be used for, and Australian Privacy Principle 1 requires the practice's privacy policy to describe what actually happens to it. A policy describing notes kept for care, while consultations feed a supplier's model, is inaccurate on its face.
The model provider behind the scribe
An ambient scribe is a chain of systems and usually more than one company. Capture happens in the vendor's application, speech recognition may be the vendor's own or a licensed service, the drafting is commonly done by a general purpose language model reached through an API, and delivery pushes the result into clinical software or secure messaging. A statement that the vendor does not train on your data describes one company's conduct at one link in that chain.
Commercial model APIs differ by tier in exactly the way that matters here. Business and enterprise terms generally exclude customer content from training by default, consumer tiers of the same provider often do not, and inputs may still be retained for a period for abuse monitoring unless a zero retention arrangement is in place. Which tier the vendor bought, and whether the no-training and retention positions are contractual or only a configuration setting someone can change, are short questions with short answers.
Where any stage sits overseas, the practice has made a cross-border disclosure, and Australian Privacy Principle 8 keeps the disclosing entity accountable for the overseas recipient's handling in most circumstances. Ask for the current subprocessor list with the country each one processes in, and for notice before it changes. A supplier that cannot produce that list has not worked through its own supply chain, which is a finding in itself.
Defaults, toggles, and what a setting cannot undo
Where a vendor offers a choice about training, read the default and the scope before you read the toggle. An opt-out that starts switched on means every consultation captured between signup and the day someone finds the setting has already been contributed. A control that lives in each user's profile instead of in tenant policy means every clinician who joins the practice starts at the vendor's default again. A control that governs transcripts but not quality-review samples covers less than it appears to.
Some suppliers offer the reverse, a model tuned on your practice's own notes so drafts come back in your house style. That is training, done deliberately, and it changes what has to be checked. Ask whether the tuned weights stay inside your tenant, whether anything learned from your notes can surface in another customer's drafts, what happens to the tuned model when you leave, and whether the underlying model provider sees the tuning data. A practice can reasonably decide that its own data should improve its own output. It cannot make the same decision on a patient's behalf about a product sold to every other practice.
Training decisions do not work symmetrically in time: turning training off today governs data from today. A model whose parameters were already fitted using your consultations does not forget when the source records are deleted, because the influence sits in the parameters and not in the row. Deleting the record is straightforward, and removing its contribution requires retraining the model without it, which no supplier does at a customer's request. So the backward-looking question belongs in writing alongside the forward-looking commitment: has any data from this practice already been used for training, fine-tuning, evaluation or human review, and if so, what data and when.
The wording to ask for in writing
Ask for the commitment in the agreement itself. Public pages change without notice and bind nobody, a questionnaire answer is a statement by a person who may not be at the company next year, and a product setting can change in any release. Where a supplier will say something in a sales meeting but will not put it in the contract, treat the contract position as the real one.
Pair each commitment with the artefact that lets you check it later. Access logs on request, a maintained subprocessor list, retention periods stated per data type, and a deletion certificate at the end of the relationship are all things a practice can obtain. Without an artefact, there is no way to test the commitment while the agreement is running.
Most of the following is wording a straightforward supplier signs without much argument. Adapt it to the agreement you are offered, and make sure each item lands in the agreement itself:
- Customer data, including consultation audio, transcripts, drafted notes and any derived artefacts, will not be used to train, fine-tune, adapt, evaluate or benchmark any machine learning model, whether the supplier's or a third party's, and whether the data is identified, pseudonymised, de-identified or aggregated.
- The supplier processes customer data only on the practice's documented instructions and for no purpose of its own.
- No subprocessor may use customer data for any of those purposes. The supplier maintains a current list of subprocessors and the country each processes in, and notifies the practice before that list changes.
- Human access to customer data is limited to named purposes, is logged, and the log is available to the practice on request.
- Retention periods are stated separately for audio, transcripts and finished notes, with the deletion method and the treatment of backups.
- On termination, all customer data and derived artefacts are deleted within a stated period, and the supplier certifies the deletion in writing.
- These terms cannot be varied by a change to the product, an online policy or a default setting without the practice's written agreement.
How aurii answers this
This section is about our product. Everything above is not.
aurii processes a consultation to produce that clinician's documentation. Health data is hosted in Australia on Azure, and each practice's data sits inside its own tenant boundary instead of a pool shared across customers. Access to health data is written to tamper-evident audit trails, so who reached what is recorded and can be produced on request.
Capture happens with the patient's consent, and that consent covers recording the encounter so the clinician can produce an accurate note. The clinician then reviews, edits and signs every note, and nothing enters a patient record without that step. The clinical record stays the practice's from capture to signature.
Put the requirement in the agreement instead of accepting it from an article. Ask aurii for the clauses above in writing, ask every other tool you are assessing for the same clauses, and compare who signs them and who negotiates them down. Where a supplier has already settled the question internally, the wording is not difficult.
Common questions
It depends entirely on the supplier and on what the practice signed, so the answer sits in the contract and not in the category. Some ambient scribes process consultations only to produce that practice's documentation, and some retain audio and transcripts as material for improving models. Ask what copies exist after the note is signed, who and what can read them, for which purposes and for how long, and require the answer in the agreement.
It is if the individual remains reasonably identifiable, which is a contextual test rather than a box to tick. Removing names and formal identifiers from a transcript does nothing about voice, rare diagnoses, dates, place names and relationships, all of which support re-identification when combined with other data. Treat any claim of de-identified training data as a technical claim needing evidence, including who assessed the residual risk and against what release context.
Not by deleting the records. Once a model's parameters have been fitted using your consultations, the influence sits in those parameters, and removing the source data does not remove it. The only reliable remedy is retraining without the data, which suppliers do not do at a customer's request. So the backward-looking question, whether any of your data has already been used, belongs in writing at the same time as the forward-looking commitment.
That is a separate question from what the scribe vendor does and it needs its own answer. Most scribes send content to a general purpose model through an API, and the training and retention terms of that API depend on the tier the vendor bought. Ask which providers sit in the chain, in which countries they process, and whether the no-training and retention positions are contractual or only a setting someone can change.
No. The practice is the entity that collected the health information, and any secondary use has to satisfy Australian Privacy Principle 6, which for sensitive information means consent or a use the patient would reasonably expect that is directly related to their care. A term accepted by a practice manager at signup is not consent from the patient, whose agreement was to being recorded so the clinician could write an accurate note.
Yes. Australian Privacy Principle 5 requires patients to be told at collection what their information will be used for and who it goes to, and Australian Privacy Principle 1 requires the practice's privacy policy to reflect what actually happens. The simpler course is to remove the question by contracting the use away, which keeps the conversation in the room about recording for the note and nothing wider.
This is general information about privacy obligations and vendor arrangements in Australian practice. It is not clinical or legal advice. More guides sit on the resources hub. If your practice needs a question answered before it adopts AI documentation, tell us and we will write it: hello@aurii.com.au.