Enhancing diagnostic accuracy in rare and common fundus diseases with a knowledge-rich vision-language model
AI Generated Summary*
RetiZero, a fundus vision-language foundation model, was pre-trained on 341,896 image-text pairs from 29 public datasets, 180 ophthalmic publications and online sources, spanning more than 400 retinal and optic nerve conditions. A frozen masked-autoencoder backbone is paired with low-rank adapters, contrastive image-text learning and Dirichlet-based feature calibration. Without task-specific training, Top-5 accuracy reached 0.840 across 15 categories (30,089 images) and 0.756 across 52 categories (7007 images), each ahead of FLAIR; retrieval Top-5 scores of 0.950 and 0.886 beat FLAIR and RETFound. Among 19 ophthalmologists grading 104 images, unaided accuracy ranged from 0.337 to 0.788, and 18 improved with model support, junior readers gaining most (18.4%, versus 12.3% senior and 10.8% expert). External cross-domain AUCs stayed at or above 0.912.
*This summary was generated by AI and is published unedited. Oku does not alter these summaries. It may contain errors or omissions and is provided for general informational purposes only. Oku does not guarantee its accuracy, completeness, or reliability. For authoritative information, please refer to the original, peer-reviewed article.