Explainability has been the standard institutional answer to clinician distrust of AI: show the reasoning and the user can evaluate it. An experiment involving 214 physicians suggests that the mechanism does not work the way the answer assumes.

Participants reviewed cases with a model recommendation, half with a short natural-language explanation and half without. Explanations raised agreement with the model by nineteen percentage points. The effect was essentially identical for cases where the model was correct and cases where it was wrong, including cases where the explanation cited a finding that was not in the chart.

The authors describe this as a fluency effect rather than a reasoning effect: a coherent explanation increases confidence in the conclusion without being independently verified. Several participants in the debrief said they had scanned the explanation for obvious errors rather than checking it against the record, which under time pressure is a rational allocation.

The uncomfortable implication is that explanation quality and explanation safety may pull in opposite directions. The authors suggest pairing explanations with an explicit verification prompt — a required click confirming the cited finding — and note that clinicians in a pilot found this deeply annoying and were measurably more accurate.