Evidence landscape: Ambient clinical documentation AI
What the published evidence record shows — and does not show — for AI tools that listen to a clinical encounter and produce a draft note. This page names no vendors and reviews no products.
What this page covers
Ambient clinical documentation AI refers to tools that use speech recognition and large language models to listen to a patient-clinician encounter, in real time or from a recording, and produce a structured or semi-structured clinical note for the clinician to review and sign. The category is sometimes called "AI scribes" in marketing, though the term is imprecise and covers a range of architectures from extractive summarisation to generative drafting.
This page describes the state of independent evidence, the regulatory landscape, the transparency and security norms, and the questions a buyer should ask — drawn from peer-reviewed literature, regulatory databases and published frameworks. It is not a product review, and it makes no claim about any specific tool.
The Institute makes no clinical, safety, regulatory or outcome claim about any named product.
The state of independent validation
The central problem a buyer faces in this category is that the published evidence base is overwhelmingly vendor-originated. Most validation studies are funded by the vendor, conducted at a single site, in a single specialty, with sample sizes that are too small to detect clinically meaningful error rates.
The methodological issues recur across the literature:
- Pre-post designs without controls. The most common study design compares documentation time before and after deployment. Without a control group, secular trends — clinician learning curves, EHR updates, staffing changes — cannot be separated from the tool's effect.
- Single-specialty, single-site samples. A tool validated in outpatient primary care encounters tells a buyer very little about its performance in an emergency department, a psychiatric intake or a surgical pre-operative consultation, where the audio environment, the clinical complexity and the documentation standard all differ.
- Self-reported outcome measures. Many studies rely on clinician satisfaction surveys rather than independent audit of note accuracy, completeness or clinical fidelity. Satisfaction is a legitimate outcome, but it is not a proxy for whether the note is correct.
- Short follow-up periods. Most published studies cover weeks to a few months. The question of whether time savings persist once the novelty effect fades, or whether editing burden increases as clinicians learn where the tool makes characteristic errors, is largely unaddressed in the literature.
- Vendor-selected populations. When the vendor selects the site, the specialty and the clinicians, the resulting sample is unlikely to represent the deployment conditions a health system will actually face.
Independent, peer-reviewed studies — conducted by researchers with no commercial relationship to the vendor, on a population the vendor did not select — remain rare in this category. This does not mean the tools do not work. It means the evidence a buyer can show a board, a regulator or a malpractice carrier is thinner than the marketing suggests.
The gap that matters most is accuracy measurement. For a tool that generates clinical documentation, the critical question is not how much time it saves but how often it fabricates clinical detail that was not in the encounter, omits detail that was, or introduces a factual error in a medication, dose or lab value. That error rate, stratified by specialty, encounter complexity and audio quality, is the number a buyer needs — and the number the published literature almost never provides with independent verification.
The regulatory picture
Most ambient clinical documentation tools currently operate outside FDA device regulation. This is a regulatory position, not a regulatory exemption, and understanding the distinction matters for procurement.
The FDA maintains a public list of AI/ML-enabled medical devices that have received marketing authorisation (FDA, "Artificial Intelligence and Machine Learning (AI/ML)-Enabled Medical Devices," updated quarterly). The majority of devices on that list are in radiology, cardiology and other imaging-adjacent specialties. Clinical documentation tools are largely absent from the list.
The reasoning is generally that a documentation tool assists with administrative workflow rather than making or materially supporting a clinical decision. Under the 21st Century Cures Act, certain clinical decision support functions are excluded from the definition of a medical device when they meet specific criteria, including that the clinician independently reviews the basis for the recommendation (21st Century Cures Act, Section 3060, enacted December 2016). A tool that drafts a note for a clinician to review and sign has a plausible argument that it sits within this framework.
The ambiguity is real and it is worth naming. A documentation tool that suggests a diagnosis code, proposes an assessment, or pre-fills a problem list is doing something closer to clinical decision support than pure transcription. Where that line sits is a matter of ongoing regulatory interpretation, and buyers should not assume that a vendor's regulatory position has been tested by the FDA merely because the vendor has not received an enforcement action.
The NIST AI Risk Management Framework, published in January 2023 (NIST AI 100-1, 26 January 2023), provides a voluntary framework for managing AI risk that applies regardless of whether a tool is FDA-regulated. The Coalition for Health AI (CHAI), based at Duke University, has published an assurance framework intended to provide a governance structure for health AI tools including those outside FDA jurisdiction (CHAI, "Blueprint for Trustworthy AI Implementation Guidance and Assurance for Healthcare," version 1.0, April 2023). Both frameworks are relevant to procurement even when the tool is not a regulated device.
The transparency gap
Transparency in this category is poor by any published standard. The concept of a model card — a structured disclosure of a model's intended use, training data, evaluation metrics and known limitations — was formalised by Mitchell et al. in 2019 (Mitchell et al., "Model Cards for Model Reporting," FAT* Conference, January 2019). In this product category, published model cards are rare.
The consequences for a buyer are concrete:
- Training data disclosure is almost nonexistent. A buyer typically cannot determine what data the model was trained on, whether it included data from their patient population or a meaningfully similar one, or whether it was trained on clinical audio at all versus general-purpose speech and text data. This makes it impossible to assess generalisability without conducting a local pilot.
- Versioning policies are uncommon. When a vendor updates the underlying model, a buyer rarely learns what changed, when it changed, or whether the performance characteristics that informed the original procurement decision still hold. A tool evaluated in Q1 may be running a different model by Q3 with no notification.
- Subgroup performance is unreported. Even where a vendor publishes aggregate accuracy figures, performance stratified by patient demographics, clinician accent, medical specialty, or encounter complexity is almost never disclosed. A tool that performs well on native-English-speaker primary care encounters may perform very differently on a non-native-speaker psychiatric evaluation — and the buyer cannot tell from the published materials.
- Failure modes are undocumented. Known limitations, edge cases and characteristic error patterns are almost never published. A clinician using the tool cannot look up the situations in which it is known to fail, and an organisation deploying it cannot build monitoring around documented failure modes because none have been documented.
The ONC's Health Data, Technology, and Interoperability (HTI-1) final rule, published in December 2023 (ONC, HTI-1 Final Rule, 88 FR 85824, 13 December 2023), introduced transparency requirements for AI and predictive models used in certified health IT. The scope and applicability of those requirements to ambient documentation tools specifically depends on how the tool integrates with the certified EHR, and this is an area where the regulatory picture continues to develop.
The security question
For a tool that processes real-time audio of clinical encounters, the security surface is larger than for most health IT products. The data in question is not only protected health information under HIPAA but includes the unfiltered content of a clinical conversation — which may contain information the patient did not intend to include in their medical record.
SOC 2 Type II attestation and a signed Business Associate Agreement are table stakes, not differentiators. The questions that matter for this category are more specific:
- Data training rights. Does the vendor use customer audio or generated notes to train or fine-tune shared models? If so, is that use default-on (opt-out) or default-off (opt-in)? The difference matters enormously. A vendor that uses patient encounter audio to improve a model serving other customers is doing something the patient almost certainly did not consent to, regardless of what the BAA permits. Policies vary widely across the category.
- Audio retention. How long is the raw audio retained, where is it stored, and who can access it? A generated note that enters the EHR is governed by the health system's retention policies. The source audio that produced it may be governed only by the vendor's policies, which may not align.
- Processing location. Whether audio is processed on-device, at the edge, or in a cloud environment affects the data exposure surface. A vendor asserting "HIPAA compliant" without disclosing the processing architecture is not giving a buyer enough information to assess the risk.
- Third-party model providers. If the vendor relies on a third-party large language model, the buyer's data may pass through infrastructure the vendor does not control and the BAA may not cover. The subprocessor chain matters, and it is rarely disclosed before contracting.
What buyers should ask
These questions are derived from the RUAIH Vendor Score method and the due diligence checklist, adapted for this product category. They are designed to be put to a vendor in writing, with the expectation of a written answer.
On validation
- Has any peer-reviewed, independently conducted study validated this tool's accuracy? If so, provide the citation, the study population, and whether any author had a commercial relationship with you.
- What is the measured rate at which the tool fabricates clinical detail not present in the source encounter? Provide the methodology, the denominator, and the specialty.
- What is the measured rate at which clinically significant detail present in the encounter is omitted from the generated note?
- How was "accuracy" defined in any study you cite, and does that definition include clinical fidelity — not merely word overlap or structural completeness?
On regulatory position
- Is this product registered, cleared, or approved by the FDA? If not, what is your basis for concluding it does not require marketing authorisation?
- If the tool suggests diagnosis codes, assessments, or problem list entries, how do you distinguish that function from clinical decision support as defined under the 21st Century Cures Act?
On transparency
- Do you publish a model card or equivalent disclosure? If so, provide it.
- What is your versioning and notification policy when the underlying model changes?
- Do you report performance stratified by patient demographics, clinician accent, medical specialty, and encounter complexity? If so, provide the data.
- What are the documented failure modes and edge cases?
On security and data rights
- Is customer encounter audio used to train or fine-tune any model, whether customer-specific or shared? Is that default-on or default-off?
- How long is raw audio retained after note generation, and where is it stored?
- Does any patient data pass through a third-party model provider? If so, name them and describe the contractual protections.
- Provide your current SOC 2 Type II report, BAA, and a plain- language summary of your data training policy.
On governance
- Does the tool produce audit logs sufficient for a Joint Commission or equivalent review?
- How does the tool detect and report model drift in production?
- What is your process for notifying customers of a material change to model performance?
A vendor that answers these questions fully, in writing, with data and citations, is giving a buyer what they need to make a defensible procurement decision. A vendor that declines is not necessarily hiding something — but a buyer cannot distinguish the two from the outside, and a board will not try.
Sources
- FDA, "Artificial Intelligence and Machine Learning (AI/ML)- Enabled Medical Devices," updated quarterly. fda.gov
- 21st Century Cures Act, Section 3060, "Clarifying Medical Software Regulation," enacted 13 December 2016.
- NIST, "Artificial Intelligence Risk Management Framework (AI RMF 1.0)," NIST AI 100-1, 26 January 2023. nist.gov
- Coalition for Health AI (CHAI), "Blueprint for Trustworthy AI Implementation Guidance and Assurance for Healthcare," version 1.0, April 2023.
- Mitchell et al., "Model Cards for Model Reporting," ACM Conference on Fairness, Accountability, and Transparency (FAT*), January 2019.
- ONC, "Health Data, Technology, and Interoperability: Certification Program Updates, Algorithm Transparency, and Information Sharing (HTI-1) Final Rule," 88 FR 85824, 13 December 2023.