Foundation models are increasingly being positioned as a route toward more generalizable ophthalmic AI. But “generalizable” depends heavily on where – and on whom – a model is tested. A new study published in Science Bulletin highlights a largely overlooked source of bias: geography, and, specifically, altitude in a given region.
Researchers in Hong Kong evaluated four established foundation models – RETFound, VisionFM, FLAIR, and CLIP – using color fundus photographs collected from medical centers spanning altitudes of approximately 40 ft to more than 12,000 ft. The cohorts included patients from plain, high-altitude, and very-high-altitude regions, with three additional high-altitude centers used for external validation. The team also developed MIXFound, an ensemble framework that combines representations from multiple frozen foundation models using performance-informed weighting and trainable classifier heads.
The results exposed a clear altitude-associated performance gap in the existing foundation models. Across the datasets, every individual model showed some deterioration as altitude increased, with cross-altitude differences exceeding 14 percentage points in AUROC for some models. The team’s newly developed MIXFound proved to be more resilient. In the internal high-altitude cohort, it achieved an AUROC of 85.56 percent, while in the very-high-altitude cohort it maintained an AUROC of 81.23 percent. External testing produced similarly encouraging results, including an AUROC of 90.02 percent in one very-high-altitude cohort.
The ensemble’s advantage was also statistically robust: MIXFound significantly outperformed the four comparator models in 27 of 28 cohort-level comparisons after Bonferroni correction. It also performed better than simpler ensemble strategies, including uniform averaging, temperature-scaled averaging, logistic regression stacking, and a gating-network approach.
So why might altitude matter for these models? Analysis of healthy eyes found significant differences in 12 of 16 retinal vascular parameters across four altitude groups. The study authors point to features such as increased vessel density, arterial dilation, and caliber remodeling as a possible anatomical contribution to the observed domain shift – although they do stress that these findings are hypothesis-generating rather than evidence of a causal effect.
There are also several caveats included in the study: all cohorts came from Chinese medical centers, and altitude itself may be acting as a proxy for a combination of environmental, demographic, imaging-device, referral, and acquisition differences. Residual age- and sex-related disparities also remained after applying MIXFound.
The broader message is difficult to ignore: strong performance at sea level does not guarantee reliable deployment elsewhere. As ophthalmic AI moves toward wider clinical use, geographic diversity may need to become as routine a part of validation as demographic diversity.