The Operations Edge Your downtime plan does not cover degradation
The Operations Edge · Issue 7 · August 20, 2026

Your downtime plan does not cover degradation

Your continuity plan detects outage rather than degradation, so a tool that stays up and quietly gets worse trips none of the machinery you already own and needs its own signal, its own declaring authority and a manual throughput number somebody has actually measured.

The tell is a meeting. Somebody who works a queue — coding, the inbox pool, referrals — says the output has felt off for a few weeks. The room’s first question is whether the tool is down. It is not. The status page is green, the interface loads, the integration is passing its checks, and the queue is moving at roughly the rate it moved last month. So the item becomes a request for data, the data lives in a report nobody owns, and the tool keeps running on every case until the committee next sits.

That is not a technology failure. It is a category failure. You have excellent machinery for a tool that stops and almost none for a tool that gets worse.

Downtime procedures are built around a binary. Something is available or it is not, and the whole apparatus follows from that: a monitor that goes red, a threshold, a named person who declares it, a paper packet, a drill you run twice a year. Every part of that chain is triggered by absence. An AI tool in production rarely gives you absence. It gives you a slow change in the quality of what it hands to a human, which is invisible to availability monitoring by construction, because nothing is unavailable.

Four things follow, and each of them is an operations problem rather than a technical one.

Detection is a person, not a monitor. An outage announces itself to the system. Degradation appears as a pattern across cases, where no single case looks wrong enough to escalate on its own. The people positioned to see the pattern first are the ones working the queue all day, which means the signal arrives as “it feels off lately” from a coder or a scheduler. That sentence loses every argument it is put into, because it is up against a dashboard that is green and a vendor whose service level is being met. So it gets raised once, informally, and not again.

Nobody can declare it. Downtime has a declaring authority and a written threshold. Degradation usually has neither, so the decision to pause routes to whichever body governs the tool, and that body meets monthly. Between meetings the default is to keep using it on every case. The default is the decision, and nobody experiences making it.

The manual baseline has quietly expired. This is the expensive one. When you deployed, you resized around the assisted throughput. The queue target came down, two positions went unfilled through ordinary attrition, the per-case time in the staffing model was revised, and a step somebody used to do by hand stopped being trained on for new starters. “Revert to manual” was a true statement on the day you signed the contract and stopped being true some quarter afterwards. What is missing is not the capability. It is the number: what your team can actually clear per day, unassisted, at today’s headcount and today’s skill mix. Almost nobody has remeasured it, which means the fallback in the plan is a memory of an organization that no longer exists.

Your agreement may not define this as an incident at all. Uptime is measurable and contractual. Output quality drifting is neither, unless somebody defined it when the contract was written. Two questions worth putting to whoever holds the vendor relationship: what does this agreement name as an incident, and what happens when the tool is up and the output has changed. Those are legal and contracting questions with real answers in your specific paperwork, and they are not answerable from the outside.

The frameworks people cite in this area already expect the capability, which is worth knowing before somebody tells you it is a novel ask. The NIST AI Risk Management Framework, published by NIST in January 2023, puts mechanisms for disengaging or deactivating a system whose performance is inconsistent with its intended use inside its MANAGE function, alongside incident response. And the HTI-1 final rule, published by ONC in December 2023, requires certified decision support to surface source attributes describing how an intervention was developed and validated. Read those together and the gap is visible: the second gives you disclosure at the point of purchase, and the first asks for something you operate every day. A description of how a tool was built is not a measurement of how it is running this week.

I am not going to give you a rate at which deployed tools drift. I have not seen a figure I would stand behind for a general estimate, and a plausible invented one would be worse than none, particularly in a piece arguing that you should go and measure something. The argument does not need the number. It is structural: the detection path, the declaring authority, and the manual figure. You can audit all three by Monday afternoon without opening a project.

Pick one deployed tool, the one furthest into production rather than the newest, and run three checks.

  1. Who noticed last? Ask two people who work its output daily whether what the tool hands them has changed in the last month. If the answer contains “sometimes” or “lately”, you have a live signal with no channel, and it has probably been live for a while.
  2. Who can turn it off? Name the person, in one sentence, without a conditional. If naming them takes a paragraph or ends in a committee, you do not have a declaring authority, you have an agenda item.
  3. What is the manual number? Ask for current unassisted throughput per person per day and the date it was last measured. If that date is older than the deployment, your continuity plan is describing a team you used to have.

Then do one thing. Write the degradation entry for that tool, on one page: the signal somebody watches, the threshold that trips it, the named human who can declare it, the measured manual throughput with the date it was measured, and who gets told inside the hour. File it where the downtime procedure lives, not in the AI governance minutes, because the person who needs it at seven on a Tuesday morning only ever looks in one of those two places.

None of this requires new software, a vendor conversation, or budget. It requires deciding that quality is an availability question and giving somebody the authority to say so out loud. A tool that stops is an incident that runs itself. A tool that gets worse is a decision you are making every day by not making it.

One operational argument a week

The Operations Edge lands each Monday: a hook, one thing to use before lunch, and the full argument here in the archive. No vendor sponsorship, ever.

The instruments behind the writing

Every framework in the series is published as a working file: registers, protocols, audit rubrics and unit-economics models, sized to be used rather than admired.

See the toolkits The library

← Ambient AI reaches the nursing station The quiet period is the cheap period →

Published under the Institute's editorial standard.

Author: Neel Chauhan, MD MBA, physician-executive and founder of the Healthcare AI Institute. Last reviewed against the standard on 2026-08-17.

Drafted as issue 7 of The Operations Edge, argued from operating experience rather than from a dataset. Quantitative claims are omitted where no dated, named source was available; the two framework references are cited by publisher and date.

The Institute accepts no vendor sponsorship, holds no vendor equity and takes no referral fees.