Draft generation for patient portal messages is among the most widely adopted clinical AI features, and the internal audits have generally focused on clinical accuracy, where results are good. A patient-side evaluation at one system looked at something else.
Patients rated generated-and-approved replies as clear and complete, and rated them lower than human-written replies on whether the responder had understood their concern. The free-text comments were specific: the replies answered the question asked and did not acknowledge the worry behind it.
This is a solvable problem and the fix has been mostly editorial. The system revised its drafting prompt to open with an acknowledgment of the patient's stated concern in their own terms, and required the responding clinician to personalize at least one sentence before sending.
Ratings recovered to near parity with human-written replies. The clinicians involved reported that the personalization requirement added roughly eight seconds per message, which the system considered an acceptable price and which the audit's authors note would have been invisible to an accuracy-only review.