The practice of auditing clinical models for subgroup performance disparities has spread quickly, driven by a mix of accreditation pressure, internal policy and genuine concern. Survey data now suggests the audits are being run and rarely acted upon.

Sixty-three percent of responding systems reported conducting subgroup performance audits on at least one deployed model in the past year. Eighteen percent reported making any change as a result — a threshold adjustment, a scope restriction, a retraining, a decommissioning.

The reasons given by respondents were mostly practical rather than dismissive. Audits frequently find disparities that are small, statistically fragile at subgroup sample sizes, or attributable to differences in outcome ascertainment rather than the model. Distinguishing these requires analytic capacity that most audit programs do not have.

The recommendation emerging from the survey authors is to fund the second half of the work. An audit function without a remediation pathway generates documentation of known problems, which is arguably a worse position than not looking — both ethically and, several respondents noted, in discovery.