How we score healthcare AI vendors
A method you cannot inspect is not a method. This is the whole of it, published so a vendor, a buyer or a critic can argue with it before anyone relies on a number.
What this scores, and what it does not
The RUAIH Vendor Score measures the quality of the evidence a buyer can independently verify about a healthcare AI product. It does not measure whether the product works.
That distinction is the entire instrument, and it is not a hedge. Nobody outside a vendor can honestly score clinical performance from the outside: you would need the model, the data, the deployment and a study. What a buyer can establish, before signing anything, is whether the vendor has published enough for a reasonable person to evaluate the claim. A product with excellent real-world performance and no published evidence scores badly here, and that is correct, because at procurement time a buyer cannot tell the two apart.
So the score answers one question. If you had to defend this purchase to your board, your regulator or your malpractice carrier, what could you actually show them?
The Institute makes no clinical, safety, regulatory or outcome claim about any named product. We score the evidence about the claim, never the claim.
Why it uses the same scale as your own readiness score
The RUAIH readiness score already grades a health system's own controls on four levels, where the distinction between documented and operating carries the finding. The vendor score deliberately mirrors that shape.
The result is one method pointed in both directions. Assess your own readiness, assess the tool you are about to buy, and compare them dimension by dimension. A tool at level 3 on every dimension dropped into an organisation full of ones is not a safe deployment, it is an expensive one. That comparison is the reason to run both.
The scale
| Level | Name | What it means |
|---|---|---|
| 0 | Absent | Nothing on the public record. |
| 1 | Asserted | The vendor states it. No artefact a third party can inspect. |
| 2 | Documented | A public artefact exists and says what the vendor says it says. |
| 3 | Verified | An independent party confirmed it, and the confirmation is dated and citable. |
The gap between 1 and 2 is where most vendor marketing sits. The gap between 2 and 3 is where most procurement risk sits.
The six dimensions
1. Independent validation
| 0 | No published validation of any kind. |
| 1 | Performance figures in marketing materials, with no methods, population or denominator. |
| 2 | A vendor-published study or technical report stating methods, population and limitations, including a preprint. |
| 3 | Peer-reviewed, or conducted by a party with no commercial interest, on a named population, with methods a reader can inspect. |
2. Regulatory position
| 0 | Regulatory status cannot be determined from public materials. |
| 1 | A status is asserted with no citation, or asserted in wording a reasonable reader could take to mean more than it does. "FDA registered" is the common example, and it is not a clearance. |
| 2 | Status is stated and consistent with public FDA records, or the product sits outside device regulation and the vendor explains correctly why. |
| 3 | Clearance, De Novo or approval is on the public record with a citable number and date, and the marketed indication matches the cleared indication. |
3. Transparency
| 0 | No model documentation of any kind. |
| 1 | A marketing-level description of the model. |
| 2 | A published model card or equivalent covering intended use, a description of training data, and known limitations. |
| 3 | All of the above, plus performance reported by subgroup, plus a stated versioning policy so a buyer knows when the thing they evaluated has changed. |
4. Security and data rights
| 0 | Nothing published. |
| 1 | Security is asserted. No certification, no terms available before contracting. |
| 2 | A current third-party attestation such as SOC 2 Type II or HITRUST, and a BAA or DPA available for inspection. |
| 3 | Both of the above, plus explicit plain-language terms that customer data is not used to train shared models, with any such use default-off rather than opt-out. |
5. Governance artefacts
| 0 | None. |
| 1 | Audit logging is mentioned somewhere, unspecified. |
| 2 | The product ships audit logs, captures human overrides, and provides a monitoring view. |
| 3 | All of the above, plus evidence exportable in a form suitable for a Joint Commission or CHAI-aligned review, plus a documented process for detecting model drift. |
6. Commercial terms
| 0 | Nothing public. Everything sits behind an NDA. |
| 1 | The pricing model is described in general terms. |
| 2 | Public pricing, or a published standard agreement with visible exit rights. |
| 3 | Public pricing plus data portability on exit, capped price escalation, and a service level with actual remedies. |
Governance artefacts is the dimension nobody else scores, and it is the one that decides whether you can survive an audit two years after go-live.
The scoring rules
Report each dimension, not a composite. We do not publish a composite number or an overall grade. Each dimension is reported as a finding — a statement you can check for yourself in a few minutes, and correct us on if we are wrong. A composite is a judgement produced by weighting those findings, and it is the one thing on the page that cannot be checked, only disputed. Every published review leads with the weakest dimension. A tool at level 3 on five dimensions with a zero in security and data rights is not a strong product with a small gap, it is a procurement problem.
Absence scores zero, and we print it. "No independent validation published as of this date" is a finding. It is not a gap in our research, and we never soften it into "limited public information available".
No citation, no point. Every point above zero carries a named, dated, public source. If we cannot link it, the vendor does not get it. That rule is what makes a score defensible when it is disputed, and it is not waived for anyone.
Scores expire. Every review is re-scored at least every six months, and immediately on a material event: a clearance, a breach, an acquisition, a pricing change or a major model release. Every page carries its last-reviewed date.
A score is never negotiated. Vendors get 48 hours before publication to correct factual errors. The number is not a fact they own. Our full review policy sets out what that window does and does not cover.
Where this sits against CHAI and NIST
A score is only useful if it survives the meeting it is taken into. Most health systems already run AI governance against the NIST AI Risk Management Framework, and increasingly against the assurance work published by the Coalition for Health AI. So each dimension is mapped onto both, and the mapping is published rather than asserted.
The Institute is not accredited, endorsed, certified or reviewed by CHAI, NIST, the Joint Commission or anyone else. This is our method, mapped onto their published frameworks so that a score arrives in language your board already uses. Nothing more than that should be read into it.
| Dimension | What it speaks to in CHAI's work | NIST AI RMF | The question your board is really asking |
|---|---|---|---|
| Independent validation | Efficacy, and local validation before use | MEASURE | Has anyone outside the vendor shown this works, and on which population? |
| Regulatory position | Safety | MAP | Is this a regulated device, and does the marketed indication match the cleared one? |
| Transparency | Usability, at the ethical design and engineering stage | MAP | Do we know what it was trained on, what it is for, and when it last changed? |
| Security and data rights | Security and privacy | GOVERN | Where does our patient data go, and can it be used to train a shared model? |
| Governance artefacts | Deployment and monitoring | MANAGE | Two years after go-live, can we produce evidence that this operated as intended? |
| Commercial terms | No direct CHAI analogue. This is procurement rather than assurance. | GOVERN | If this fails, or the vendor is acquired, can we leave and take our data? |
The four NIST functions, in their own terms:
| GOVERN | Cultivates and implements a culture of risk management across the organisation. Cross-cutting. |
| MAP | Establishes the context, and identifies what the system is for and what could go wrong. |
| MEASURE | Analyses, assesses, benchmarks and monitors the risk and its impacts. |
| MANAGE | Allocates resources to the risks that were mapped and measured, and plans the response. |
Two things are worth saying plainly about that table. Commercial terms has no CHAI analogue, because CHAI is an assurance body and exit rights are a procurement problem, and we score it anyway because a health system that cannot leave a bad contract has a risk its assurance framework will never show it. And nothing here scores MEASURE the way NIST means it, because measuring a model's behaviour requires access we do not have. What we can establish is whether anyone has measured it and published the result.
What we would test, if we could
We do not test, trial, pilot or benchmark products. That is a real limit and it is worth being concrete about what it costs, rather than leaving it as a disclaimer.
An independent evaluation with vendor access could establish things the public record cannot. For an ambient documentation tool: the rate at which clinical detail is fabricated when the audio is hard, the rate at which a lab value or a dose goes missing, and how much editing time actually survives to the signed note. For a revenue cycle tool: first-pass appeal accuracy against a standard set of complex denials. Those are the numbers a buyer wants and no vendor deck reliably provides.
Every paid dossier therefore ends with those questions written out for the specific product, in a form you can put to the vendor and ask them to answer with data. That is not the same as us having run the test, and we never write as though it were. It is the difference between a buyer asking "is it accurate?" and a buyer asking "on which audio, against which reference, measured how, and will you show us?"
Known limitations of this method
Stated here rather than left to be discovered by a critic.
It rewards disclosure, and disclosure correlates with company size. A well-funded vendor can afford SOC 2 and a peer-reviewed study. A better product from a smaller company can score lower. We think that is an acceptable bias, because a buyer genuinely does carry more risk with an unevidenced product, but it is a bias and it should be read as one.
It is a point-in-time measure of public information. A vendor that shares everything under NDA and nothing publicly is penalised. That is deliberate, since evidence you cannot show anyone else is worth less at a board meeting, but it is not the same as the evidence not existing.
It does not measure fit. The right score for the wrong workflow is still the wrong purchase. That is what the readiness score and the toolkits are for.
Six dimensions is a compression. Interoperability, implementation burden and support quality all matter, and none of them are reliably verifiable from public sources, so they are described in the dossier narrative rather than scored.