A retrospective review of eighteen clinical AI implementations that succeeded in pilot and failed at scale identifies a consistent pattern that has little to do with the technology.

During pilots, a dedicated implementation lead handled the parts of the workflow the software did not cover: chasing missing data, fixing bad routing, retraining users who drifted, and escalating when a clinician reported something odd. At scale, that role was not replicated, on the reasonable-sounding theory that the pilot had proven the workflow.

What the pilot had actually proven was that the workflow succeeded with a person compensating for it. The review's authors describe this as a systematic overestimation of readiness, and note that pilot reports almost never itemize the implementation lead's activities, because those activities are not the intervention being studied.

Their recommendation is procedural and cheap: require pilots to log every manual intervention, and treat that log as a specification for what must be automated or staffed before scaling. Three of the eighteen organizations have adopted the practice. All three report that the log was longer than anyone expected.