arXiv:2606.06407cs.CVcs.IR2026-06

用实体感知框架提升医学影像对比诊断能力

A Vision-language Framework for Comparative Reasoning in Radiology

论文配图:A Vision-language Framework for Comparative Reasoning in Radiology
图 1 · 摘自论文原文
  • 基于解剖结构和病灶构建跨图像对比推理机制
  • 在12项检索任务中召回率最高,跨中心效果提升6.0个百分点
  • 适合临床辅助诊断与纵向随访场景的AI系统开发者

医学影像人工智能在单图解读上表现优异,但在实际放射学实践中,诊断与随访依赖于对既往影像及相似病例的对比分析。本文将影像对比建模为实体感知的跨图像推理问题,提出支持参考病例检索与时间序列对比解读的框架。构建了包含超过69万张图像、来自16万余名患者、覆盖8所机构、4个国家、7种成像模态的大型对比影像资源MedReCo-DB。报告被分解为解剖结构、异常发现与病理状态,用于指导实体条件下的检索与对比视觉问答。基于此资源,开发了可控制检索临床相似病例的MedReCo实体感知视觉编码器,以及支持生成式时序变化解读的MedReCo-VLM。在内部、外部及跨中心评估中,MedReCo在全部12个内部检索任务中达到最高Recall@1,外部检索平均提升6.0个百分点;在临床易混淆差异组中持续优于最强基线。MedReCo-VLM在所有对比生成评估中表现最佳,在胸片随访中准确率提升14.5–46.5个百分点,CT提升13.0–27.9个百分点。结果表明,从常规临床数据中可大规模学习实体感知的对比推理,为医学影像AI提供更贴近临床的基础。

原文摘要 · Abstract (English)

Medical imaging artificial intelligence has achieved strong performance in isolated image interpretation, but remains poorly aligned with radiological practice, where diagnosis and follow-up rely on comparison across prior studies and analogous reference cases. Here we formulate radiological comparison as an entity-aware cross-image reasoning problem and introduce a framework that supports both reference-case retrieval and temporal comparative interpretation. We construct MedReCo-DB, a large-scale comparative imaging resource derived from routine image-report pairs, comprising more than 690,000 images from over 160,000 patients across eight institutions, four countries and seven imaging modalities. Reports are decomposed into anatomical structures, abnormal findings and pathological conditions to provide supervision for entity-conditioned retrieval and comparative visual question answering. Using this resource, we develop MedReCo, an entity-aware visual encoder for controllable retrieval of clinically analogous cases, and MedReCo-VLM, a vision--language extension for generative interpretation of interval change. Across internal, external and cross-center evaluations, MedReCo achieved the highest Recall@1 in all 12 internal retrieval settings and improved external retrieval by a mean of 6.0 percentage points. In clinically confusable differential groups, it consistently outperformed the strongest baselines. MedReCo-VLM achieved the best performance across all comparative generation evaluations and improved longitudinal follow-up accuracy by 14.5-46.5 percentage points on chest radiographs and 13.0-27.9 percentage points on CT. These findings suggest that entity-aware comparative reasoning can be learned from routine clinical data at scale and may provide a more clinically aligned foundation for medical imaging AI.

医学影像对比推理多模态临床对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。