融合眼底图、文本报告和结构化数据,提升青光眼风险预测准确率。
GlaBoost: A Multimodal Structured Framework for Glaucoma Risk Stratification
- 用梯度提升融合图像、文本和临床指标三类数据
- 在两个真实数据集上超越单一与通用多模态模型
- 可解释性强,关键预测因子符合临床共识
早期精准检测青光眼对防止不可逆视力丧失至关重要,但现有AI方法多依赖单模态输入且缺乏可解释性。本文提出GlaBoost,一种多模态梯度提升框架,整合三种互补信号进行青光眼风险预测:基于预训练卷积编码器的眼底图像嵌入、基于Transformer的文本报告神经视网膜盘边缘评估编码,以及结构化眼科学生物标志物。这些模态被融合为单一表示,并由增强版XGBoost模型分类。在两个真实标注数据集上,GlaBoost持续优于单模态及通用多模态基线。特征重要性分析表明,杯盘比、盘缘变薄和ISNT规则是主导预测因子,决策结果具临床一致性与可解释性。GlaBoost为眼科多模态决策支持提供了透明且可扩展的基础。
原文摘要 · Abstract (English)
Early and accurate glaucoma detection is critical to prevent irreversible vision loss, yet existing AI methods often rely on unimodal inputs and lack interpretability. We present GlaBoost, a multimodal gradient boosting framework that unifies three complementary signals for glaucoma risk prediction: fundus image embeddings from a pretrained convolutional encoder,free-text neuroretinal rim assessments encoded by a transformer-based language model, and structured ophthalmic biomarkers. These modalities are fused into a single representation and classified by an enhanced XGBoost model.On two real-world annotated datasets, GlaBoost consistently outperforms unimodal and generic multimodal baselines. Feature importance analysis highlights the cup-to-disc ratio, rim thinning, and the ISNT rule as the dominant predictors, yielding clinically consistent and interpretable decisions. GlaBoost offers a transparent and scalable foundation for multimodal decision support in ophthalmology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。