用多模态蒸馏提升眼底图像与临床数据融合,增强心脏病筛查准确率
Cross-Modal Iteration Distillation for Robust IHD Screening: The IDNet Framework and A New Benchmark

- 设计可学习查询的跨模态蒸馏聚合器,逐步融合双眼图像与临床数据
- 在5万张眼底图上验证,模型性能优于纯图像或纯临床数据方法
- 开源英国生物银行基准数据集,支持可复现研究
彩色眼底摄影(CFP)为缺血性心脏病(IHD)筛查提供了低成本、无创途径,但现有研究受限于公开基准数据稀缺以及眼底图像与稀疏临床变量融合效果差。本文提出IDNet多模态框架,采用可学习查询的跨模态蒸馏聚合器(CDA),依次整合左眼、右眼及临床特征,缓解高维视觉特征与低维表格输入之间的不平衡问题。同时构建了可复现的英国生物银行基准,包含开源的数据清洗与质量控制流程,共涵盖50,410张图像,来自25,205名受试者。在该基准上,IDNet超越仅使用图像、仅使用临床数据及多种多模态基线模型,且CDA作为即插即用模块持续提升多个视觉编码器性能。
原文摘要 · Abstract (English)
Color Fundus Photography (CFP) offers a low-cost and non-invasive route for ischemic heart disease (IHD) screening, but current studies are limited by scarce public benchmarks and ineffective fusion of retinal images with sparse clinical variables. We propose IDNet, a multimodal framework with a Cross-Modal Distillation Aggregator (CDA) that uses learnable queries to sequentially integrate left-eye, right-eye, and clinical features, mitigating the imbalance between high-dimensional visual features and low-dimensional tabular inputs. We also construct a reproducible UK Biobank benchmark with open-source curation and quality-control pipelines, yielding 50,410 images from 25,205 subjects. On this benchmark, IDNet outperforms image-only, clinical-only, and several multimodal baselines, and CDA consistently improves multiple visual encoders as a plug-in fusion module.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。