将胸部X光纵向对比与局部定位结合,提升医学报告准确性。
CheXGround: Anatomical Region Tokens for Grounded Longitudinal Chest X-ray Interpretation

- 用解剖区域令牌表示前后影像的对应部位
- 在多任务测试中,定位准确率和时序推理能力均优于基线
- 适合需要精准图像-文本对齐的临床医生和研究者
近期放射科多模态语言模型在胸片报告生成、视觉问答和时序推理方面取得显著进展。纵向胸片解读需比较序列检查以描述变化,而视觉定位则旨在将临床语言与局部图像证据关联。尽管纵向建模和视觉定位各自推动了放射科语言模型发展,但局部视觉证据如何支持纵向解读仍研究不足。我们提出CheXGround,一种基于解剖区域的纵向胸片语言模型,通过对应解剖区域表示成对检查。CheXGround从当前与历史胸片中提取解剖区域,将其编码为时序增强的感兴趣区域(ROI)令牌,并在生成时融合全局时序图像上下文。为连接区域令牌与临床文本,我们设计了时序区域-短语对齐预训练目标,使时序解剖表征与报告中的局部短语对齐。我们在单次与纵向视觉问答、纵向发现生成、时序定位问答及解剖定位任务上评估了CheXGround,结果表明其在临床语言质量、时序推理和定位精度方面均优于近期基线。结果表明,在解剖层面组织纵向证据是强大的放射科语言建模表示方式。
原文摘要 · Abstract (English)
Recent radiology multi-modal language models have made substantial progress in chest X-ray report generation, visual question answering, and temporal reasoning. While longitudinal chest X-ray interpretation compares sequential examinations to describe change, visual grounding aims to connect clinical language with localized image evidence. Although longitudinal modeling and visual grounding have each advanced radiology language models, how localized visual evidence can support longitudinal interpretation remains under-explored. We introduce CheXGround, a region-grounded longitudinal chest X-ray language model that represents paired studies through corresponding anatomical regions. CheXGround extracts anatomical regions from current and prior radiographs, encodes them as temporally enhanced Region-of-Interest (ROI) tokens, and combines them with global temporal image context during generation. To connect these region tokens with clinical text, we propose Temporal Region--Phrase Alignment, a pretraining objective that aligns temporal anatomical representations with localized report phrases. We evaluate CheXGround on single-study and longitudinal Visual Question Answering (VQA), longitudinal findings generation, temporal grounded VQA, and anatomical grounding. Across these tasks, CheXGround improves clinical language quality, temporal reasoning, and localization accuracy over recent baselines. Our results suggest that organizing longitudinal evidence at the anatomical level is a strong representation for grounded radiology language modeling. Project page: https://adonaydem.github.io/chexground-website
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。