Clearance for a clinical AI device establishes that it performed acceptably on a validation set at a point in time. What happens afterward is, for the large majority of deployed models, unmeasured.
A systematic review published this month examined 340 cleared devices and attempted to determine whether any ongoing performance monitoring was in place at deployment sites. The authors were able to confirm continuous monitoring for thirty-one of them. For the rest, monitoring was either absent, limited to uptime and utilization, or the authors could not determine what existed.
The reasons cited by sites are not mysterious. Monitoring clinical performance requires ground truth, and ground truth requires either chart review or a downstream outcome that takes months to materialize. Both are expensive. Neither is reimbursed.
Several of the review's authors have proposed that vendors be required to supply a monitoring harness as a condition of sale, shifting the cost to the party best positioned to build it once. Industry response has been cool, on the reasonable grounds that ground truth lives in the customer's chart, not the vendor's product.