通过模拟医生诊断思路,提升早产儿视网膜病变筛查的准确率与可解释性。
Context-Aware Asymmetric Ensembling for Interpretable Retinopathy of Prematurity Screening via Active Query and Vascular Attention
- 分两路处理:结构特征用动态查询定位异常区,血管拓扑用门控多实例学习识别扭曲。
- 在188名婴儿6004张图像上,大类分期F1达0.93,严重病灶检测AUC达0.996。
- 生成反事实注意力图与血管威胁图,让模型决策过程透明可追溯,适合临床部署。
早产儿视网膜病变(ROP)是导致可预防儿童失明的主要原因之一。自动化筛查面临数据稀缺及结构分期与微血管异常共存的复杂性挑战。现有深度学习模型依赖大规模私有数据集和被动多模态融合,在小规模、不平衡的公开数据集上泛化能力差。为此,本文提出上下文感知的非对称集成模型(CAA Ensemble),模拟临床推理过程,包含两个专用分支:首先,多尺度主动查询网络(MS-AQNet)作为结构专家,利用临床上下文生成动态查询向量,空间调控视觉特征提取以精确定位纤维血管嵴;其次,VascuMIL 在门控多实例学习(MIL)框架中编码血管拓扑图(VMAP),精准识别血管迂曲。一个协同元学习器整合这两路正交信号,解决多目标诊断冲突。在包含188名婴儿共6,004张图像的高度不平衡队列上,该框架在两项不同临床任务中均达到当前最优性能:宽范围ROP分期宏平均F1得分为0.93,Plus Disease检测的AUC为0.996。关键在于,系统具备“玻璃盒”透明性,通过反事实注意力热图和血管威胁图证明临床元数据主导模型视觉搜索。此外,本研究验证了架构归纳偏置可有效弥合医学AI的数据鸿沟。
原文摘要 · Abstract (English)
Retinopathy of Prematurity (ROP) is among the major causes of preventable childhood blindness. Automated screening remains challenging, primarily due to limited data availability and the complex condition involving both structural staging and microvascular abnormalities. Current deep learning models depend heavily on large private datasets and passive multimodal fusion, which commonly fail to generalize on small, imbalanced public cohorts. We thus propose the Context-Aware Asymmetric Ensemble Model (CAA Ensemble) that simulates clinical reasoning through two specialized streams. First, the Multi-Scale Active Query Network (MS-AQNet) serves as a structure specialist, utilizing clinical contexts as dynamic query vectors to spatially control visual feature extraction for localization of the fibrovascular ridge. Secondly, VascuMIL encodes Vascular Topology Maps (VMAP) within a gated Multiple Instance Learning (MIL) network to precisely identify vascular tortuosity. A synergistic meta-learner ensembles these orthogonal signals to resolve diagnostic discordance across multiple objectives. Tested on a highly imbalanced cohort of 188 infants (6,004 images), the framework attained State-of-the-Art performance on two distinct clinical tasks: achieving a Macro F1-Score of 0.93 for Broad ROP staging and an AUC of 0.996 for Plus Disease detection. Crucially, the system features `Glass Box' transparency through counterfactual attention heatmaps and vascular threat maps, proving that clinical metadata dictates the model's visual search. Additionally, this study demonstrates that architectural inductive bias can serve as an effective bridge for the medical AI data gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。