The volume of published clinical AI evaluation has grown enormously. A methodological review of 512 studies published over eighteen months finds that the growth has been almost entirely in a category of study that cannot answer the question clinicians ask.
Seventy-eight percent of the reviewed studies reported only model performance metrics on retrospective data. Fourteen percent reported a process outcome — time to result, documentation burden, alert volume. Eight percent reported a patient outcome. Of those, eleven studies were randomized.
The authors are careful not to dismiss the retrospective literature, which is necessary and appropriate at the development stage. Their argument is about the ratio, and about the tendency for retrospective performance to be cited in procurement conversations as though it were evidence of benefit.
Their proposed remedy is a reporting standard requiring every clinical AI paper to state explicitly, in the abstract, what class of evidence it provides. Several journals have expressed interest. Whether a labeling requirement changes what gets funded is a separate question the authors do not claim to answer.