让AI像医生一样对比影像,发现临床有意义的差异
RadDiff: Describing Differences in Radiology Image Sets with Natural Language
- 用医学知识增强的多模态模型进行对比分析
- 在专家验证数据集上达47%准确率,报告引导下升至50%
- 适合医学影像分析、AI可解释性研究者使用
理解两组放射科图像的差异对生成临床洞察和解释医疗AI系统至关重要。我们提出RadDiff,一种模拟放射科医生比较推理的多模态智能体系统,用于描述成对影像研究中具有临床意义的差异。RadDiff基于VisDiff的提议-排序框架,引入四项受真实诊断流程启发的创新:(1) 通过领域适配的视觉-语言模型注入医学知识;(2) 融合图像与临床报告的多模态推理;(3) 多轮迭代假设优化;(4) 针对性视觉搜索,定位并放大显著区域以捕捉细微发现。为评估RadDiff,我们构建了RadDiffBench,一个包含57对专家验证影像对的挑战性基准,配有真实差异描述。在该基准上,RadDiff取得47%准确率,报告引导下提升至50%,显著优于通用领域基线VisDiff。我们还展示了RadDiff在新冠表型对比、种族亚组分析及生存相关影像特征发现等多样化临床任务中的泛化能力。综上,RadDiff与RadDiffBench共同构成了系统揭示放射学数据中关键差异的首个方法与基准基础。
原文摘要 · Abstract (English)
Understanding how two radiology image sets differ is critical for generating clinical insights and for interpreting medical AI systems. We introduce RadDiff, a multimodal agentic system that performs radiologist-style comparative reasoning to describe clinically meaningful differences between paired radiology studies. RadDiff builds on a proposer-ranker framework from VisDiff, and incorporates four innovations inspired by real diagnostic workflows: (1) medical knowledge injection through domain-adapted vision-language models; (2) multimodal reasoning that integrates images with their clinical reports; (3) iterative hypothesis refinement across multiple reasoning rounds; and (4) targeted visual search that localizes and zooms in on salient regions to capture subtle findings. To evaluate RadDiff, we construct RadDiffBench, a challenging benchmark comprising 57 expert-validated radiology study pairs with ground-truth difference descriptions. On RadDiffBench, RadDiff achieves 47% accuracy, and 50% accuracy when guided by ground-truth reports, significantly outperforming the general-domain VisDiff baseline. We further demonstrate RadDiff's versatility across diverse clinical tasks, including COVID-19 phenotype comparison, racial subgroup analysis, and discovery of survival-related imaging features. Together, RadDiff and RadDiffBench provide the first method-and-benchmark foundation for systematically uncovering meaningful differences in radiological data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。