Skip to content
All posts

Healthcare AI · August 3, 2026 · kaddu livingstone · 9 min read

Why Clinical AI Loses Accuracy When Moved to New Hospitals

Most clinical AI models lose accuracy when moved to a new hospital, and regulatory clearance does not measure it. What the external validation evidence shows.

Why Clinical AI Loses Accuracy When Moved to New Hospitals
Share

A clinical AI model's published accuracy describes how it behaved on data from the institution that built it. It is not a forecast of how it will behave in your hospital. Across the published evidence, the most common outcome when a model meets a new institution's data is that it performs worse, often by enough to change whether the tool is worth deploying.

That matters more, not less, for hospitals in Uganda and across the region. What predicts failure is the distance between a model's development setting and its deployment setting, and that distance is widest when the model was trained in a well-resourced academic center on a different population, different equipment, and a different disease mix. Curely's position is that advanced healthcare intelligence should not be a privilege, and a model that only works where it was built is a privilege by construction.

Degradation across institutions is the rule, not the exception

The most direct evidence is a systematic review of deep learning algorithms for radiologic diagnosis (strong evidence, systematic review). Yu, Mohajer, and Eng, in Radiology: Artificial Intelligence, screened 6,018 articles and found only 83 reporting performance on an external dataset, together describing 86 algorithms.

Of those 86, 70 (81 percent) performed worse externally. Forty-two (49 percent) dropped by at least 0.05 on the unit scale, and 21 (24 percent) by at least 0.10. The median difference was 0.046 toward worse external performance, ranging from a 0.60 decrease to a 0.13 increase.

Two details matter as much as the headline. No study characteristic predicted the size of the drop, not dataset size, not the number of contributing institutions, not prevalence, not task difficulty, so a buyer cannot read a development description and infer that a model will travel well. And these were the studies that validated externally at all. Eighty-three came out of 6,018 screened, and the review cites earlier work finding roughly 6 percent of medical imaging AI publications include external validation. The 81 percent figure describes a self-selected group of careful teams, and the authors note publication bias probably makes it an underestimate.

Only 11 of 86 studies (13 percent) followed any published reporting guideline, and equipment details were almost always missing, so a hospital cannot check whether its own scanners resemble the ones the model learned from. Every included study was retrospective, describing accuracy on stored images rather than clinical effect in a live workflow.

Regulatory clearance does not settle the question

Clearance and validation are different things, and conflating them is one of the more expensive mistakes in health-tech procurement.

An analysis of 130 FDA-cleared AI devices in Nature Medicine found 97 percent were evaluated only retrospectively, none of the high-risk devices had a prospective evaluation, and most decision summaries did not report whether the device had been tested at more than one site (moderate evidence, systematic analysis of regulatory filings). A 510(k) clearance establishes substantial equivalence to an existing device. It does not establish clinical benefit, it is not FDA approval, and it carries no weight outside the United States. CE marking and national approval in East African jurisdictions are separate processes with separate evidence requirements.

A recent study in npj Digital Medicine makes the gap concrete. Chavoshi and colleagues evaluated a commercial, FDA-cleared model for intracranial hemorrhage triage across 101,944 head CT examinations from 74,142 patients in a 17-facility health system over two years, of which 6,814 exams (6.7 percent) were positive (moderate evidence, large retrospective evaluation, single health system).

The clearance summary reported sensitivity of 96.15 percent and specificity of 94.83 percent, measured on 220 cases, with no breakdown by hemorrhage size, location, or age. In deployment, sensitivity was 82.2 percent and specificity 97.6 percent. The subgroups carry the real signal. Sensitivity reached 86.2 percent for acute hemorrhages and 95.0 percent for those larger than 10 mm, but fell to 74.8 percent at 10 mm or less, 76.0 percent for single-compartment bleeds, 54.8 percent for chronic hemorrhages, and 45.5 percent for subacute ones. In outpatients, where subtle findings are more common, it was 72.2 percent.

The authors draw the conclusion that matters. The model performed well on cases a radiologist would also catch easily, and poorly on the cases human readers are most likely to miss. A tool that fails where its user fails adds less value than its pooled accuracy suggests. Note the asymmetry in sample size too. The clearance figure rested on 220 cases, the deployment figure on more than 100,000.

Models often learn the hospital rather than the disease

The mechanism is well understood. Zech and colleagues, in PLOS Medicine, trained pneumonia detection models on chest radiographs from three sources with very different pneumonia prevalence, 34.2 percent at one hospital system against 1.2 percent and 1.0 percent at two others (moderate evidence, multi-site cross-sectional study). Sorting cases by hospital system alone, with no image analysis, achieved an AUC of 0.861 on the combined dataset.

The models learned that shortcut. The best internal performance, an AUC of 0.931 from pooled data across two systems, fell to 0.815 on the third, and the networks reliably identified which hospital system, and even which department, a radiograph came from. A model can equally learn that portable radiographs are more likely to show disease, because sicker patients are the ones imaged at the bedside. That is a fact about care logistics rather than pathology, and it is why internal testing cannot substitute for external validation.

The distance a model must travel is greater in low-resource systems

Every factor known to drive degradation is amplified when a model built in a high-income academic center is deployed in a district hospital in East Africa. Prevalence differs, and prevalence is exactly the shortcut Zech's work showed models exploit. Case mix differs, with more late presentations and a different balance of infectious and non-communicable disease. Equipment differs, and the Yu review found studies rarely publish enough technical detail for a buyer to compare. Care setting differs, and the hemorrhage study found its largest drop in outpatients, closer to the first-presentation pattern common across the region.

Verification is harder too. Measuring local performance requires a labeled reference standard, which requires clinician time that is already the scarcest resource in the system. The hospitals least able to check a vendor's claims are the ones for whom those claims are least likely to hold. One gap should also be stated plainly. We are not aware of a systematic review quantifying this degradation specifically across African health systems, and most of the literature above comes from North America, Europe, and East Asia. That absence is itself a finding, and it argues for local evaluation rather than optimism.

Tuberculosis screening shows what a disciplined answer looks like

There is a working model for this problem, from a setting familiar to readers in Uganda.

In its 2021 consolidated guidelines on tuberculosis screening, the World Health Organization recommended computer-aided detection software as an alternative to human interpretation of digital chest X-rays for screening and triage in people aged 15 and older. Two qualifications travel with that recommendation and are usually dropped in vendor material. It is conditional, and WHO graded the certainty of the underlying evidence as very low. WHO also treats a single global threshold as inadequate, publishing a calibration toolkit that instructs programs to set the score threshold for their own population and context. The body that recommended the technology told implementers not to inherit someone else's numbers.

Local evidence supports that instruction. Sung and colleagues, in The Lancet Digital Health, screened people aged 15 and over around Kampala using portable digital chest X-ray with CAD software between June 2022 and March 2024 (moderate evidence, large cross-sectional diagnostic accuracy study). Among 7,219 participants with valid Xpert results, 382 (5.3 percent) were Xpert-positive, and the software achieved an estimated AUC of 0.92. At 96.1 percent specificity, estimated sensitivity was 75.0 percent using a universal threshold of 0.65 or above, against 76.9 percent using thresholds stratified by age and sex.

Two caveats belong with those numbers. Participants scoring below 0.1 were not asked for sputum, so accuracy in that group is estimated rather than measured. And the gain from stratification is roughly two percentage points of sensitivity, meaningful at population scale but modest in a single clinic. The transferable lesson is the discipline rather than the figures. Threshold selection is a local decision, and accepting a vendor default is a clinical decision made by someone who has never seen your patients.

One serious counterargument, generalizability may be the wrong target

Futoma and colleagues argued in The Lancet Digital Health that an overly broad notion of generalizability can obscure cases where a model delivers real clinical utility within a single setting, and that the field should focus on usefulness at the bedside rather than universal transferability. This is a commentary, not empirical evidence. The argument is largely right, and it strengthens rather than weakens the case for local evaluation. A model built for one health system and working well there is a legitimate design choice, but it is not a general-purpose product. A vendor should say which setting the evidence covers, and a hospital outside that setting should assume it inherits nothing.

Five questions worth asking before signing

A vendor who cannot answer these has told you something useful.

  1. What were the development and validation populations, and how do their prevalence and case mix compare to ours?
  2. Has performance been reported on data from an institution with no role in development? If not, treat the published figure as an upper bound.
  3. What is the performance breakdown by finding size, disease acuity, care setting, and imaging equipment? A single pooled sensitivity conceals the failure modes that cause harm.
  4. What is the operating threshold, who set it, on what population, and can we recalibrate it locally?
  5. What is the plan for local evaluation and ongoing monitoring, who pays for it, and what happens when a new model version ships?

What this means for how we build

The evidence supports a narrow claim and not a broader one. It strongly supports the claim that most clinical AI loses accuracy when it moves between institutions, that regulatory clearance does not measure this, and that the loss concentrates in the subtle cases where clinicians most need help. It does not support the claim that clinical AI fails to work. CAD for tuberculosis is WHO-recommended and in routine use, and the Kampala data show clinically meaningful accuracy when thresholds are set locally.

For hospitals, an external validation study in your own setting is not an academic exercise. It is the only measurement that describes the tool you are actually buying, and its cost is small next to a triage system that quietly misses small bleeds in outpatients for two years. For builders, including us, the standard is stricter. Curely describes its capabilities as design intent unless deployment data exists, and reports performance only against the setting in which it was measured. Where we have not measured, we say so. A number borrowed from somewhere else is not evidence about anywhere else.