AUROC — the number on the baseline
When a governance committee locks a baseline of “AUROC 0.87,” it is recording the model’s ranking ability on the approval-date population: given one patient who has the condition and one who does not, the probability the model scores the first higher.
Its governance virtues: it does not depend on where the alert threshold is set, so it is comparable quarter over quarter even as operational cut-offs move — the property a drift check needs.
Its honest limits: it says nothing about calibration (whether “80% risk” means 80%), nothing about performance at the specific operating threshold clinicians experience, and it can look stable overall while a subgroup gap widens. A monitoring plan built on AUROC alone is better than none and worse than one that pairs it with sensitivity and specificity at the operating point.
← All RUAIH questions, areas and terms · The complete healthcare AI governance guide · Score your readiness