A comparative evaluation run by an academic informatics group tested four commercial triage tools and one fine-tuned open weights model on the same retrospective cohort. The open model placed first on the primary metric and second on calibration, at an inference cost roughly two orders of magnitude below the commercial licenses.
It has not been deployed anywhere, and the researchers are candid that they did not expect it to be. A model is a fraction of what a hospital buys. The rest is clearance, integration, indemnification, a support contract, and an entity that carries liability if something goes wrong at 3 a.m.
The finding that matters is therefore not about open models winning. It is about how much of the price of clinical AI is attached to things other than the model, and whether that ratio is stable. Several of the health systems we spoke with are exploring a middle path: open weights, internally validated, with an external firm contracted for the regulatory and monitoring wrapper.
Nobody has completed that path yet. The first organization to do it will produce a more interesting result than the benchmark did.