Three years ago a substantial number of academic health systems were training clinical language models on their own corpora. The current census is smaller, and the systems that stopped generally give the same reason: the frontier moved faster than their retraining cycle.
The systems still building fall into a recognizable pattern. They have a use case where the data cannot leave — behavioral health notes, certain research cohorts, cases under legal hold — or a domain where general models remain weak enough that local fine-tuning is decisive. Both are narrow, and both are defended clearly by the teams running them.
What has largely disappeared is the general-purpose institutional model, trained on everything, intended to serve every use case. Several informatics directors described it, with hindsight, as a project whose main output was a data platform. That platform is now serving purchased models, which most describe as a reasonable trade.
The capability that turned out to be durable was evaluation. Systems that built their own models developed the ability to test one, and that skill transferred to procurement in a way the model weights did not.