用文本引导的细粒度表示学习,提升胸部X光片进展检测精度
CheXLearner: Text-Guided Fine-Grained Representation Learning for Progression Detection
- 基于双曲几何的结构对齐模块,精准捕捉解剖区域变化
- 在解剖区域进展检测上达到81.12%准确率,较基线提升17.2%
- 适合医学影像分析、疾病进展预测等临床场景使用
时间性医学图像分析对临床决策至关重要,但现有方法或仅粗粒度对齐图像与文本,导致语义错配;或仅依赖视觉信息,缺乏医学语义整合。我们提出CheXLearner,首个端到端框架,融合解剖区域检测、基于黎曼流形的结构对齐与细粒度区域语义引导。所提出的Med-Manifold对齐模块(Med-MAM)利用双曲几何稳健对齐解剖结构,并捕捉跨时序胸片中的病理学有意义差异。通过引入区域进展描述作为监督信号,CheXLearner增强跨模态表示学习并支持动态低层特征优化。实验表明,其在解剖区域进展检测上达到81.12%(+17.2%)平均准确率和80.32%(+11.05%)F1分数,显著优于现有最优基线,尤其在结构复杂区域表现突出。此外,模型在下游疾病分类任务中获得91.52%平均AUC,验证其优异特征表示能力。
原文摘要 · Abstract (English)
Temporal medical image analysis is essential for clinical decision-making, yet existing methods either align images and text at a coarse level - causing potential semantic mismatches - or depend solely on visual information, lacking medical semantic integration. We present CheXLearner, the first end-to-end framework that unifies anatomical region detection, Riemannian manifold-based structure alignment, and fine-grained regional semantic guidance. Our proposed Med-Manifold Alignment Module (Med-MAM) leverages hyperbolic geometry to robustly align anatomical structures and capture pathologically meaningful discrepancies across temporal chest X-rays. By introducing regional progression descriptions as supervision, CheXLearner achieves enhanced cross-modal representation learning and supports dynamic low-level feature optimization. Experiments show that CheXLearner achieves 81.12% (+17.2%) average accuracy and 80.32% (+11.05%) F1-score on anatomical region progression detection - substantially outperforming state-of-the-art baselines, especially in structurally complex regions. Additionally, our model attains a 91.52% average AUC score in downstream disease classification, validating its superior feature representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。