构建首个胸部X光片时序推理数据集,助力模型理解病灶演变。
CheXTemporal: A Dataset for Temporally-Grounded Reasoning in Chest Radiography

- 基于配对前后胸片,标注病灶时空变化与进展类型。
- 模型在稳定、缓解等细微变化上表现差,对严重恶化更敏感。
- 适合研究医学影像时序分析与跨域泛化能力的学者使用。
胸部X光解读需对前后检查进行时序推理,但现有视觉语言模型多基于静态图像-报告对训练,缺乏显式时序监督。我们提出CheXTemporal数据集,包含成对的前后胸片(CXR),并标注病灶级的时间与空间信息。数据集涵盖五类进展分类(新发、恶化、稳定、改善、消失),提供局部空间标注、跨研究显式时空对齐及多源覆盖以支持跨域评估。此外,我们构建了一个含28万对样本的银质数据集,通过自动推导获得时序与解剖监督,用于弱监督下的大规模评估。利用这些资源,我们在零样本设置下评估多个先进视觉语言模型在定位与进展分类任务上的表现。在金标准与银标准评估中,当前模型普遍存在空间定位不准、细粒度时序推理不足、分布外泛化性差等问题。尤其在稳定与缓解等微妙状态识别上表现显著落后于恶化类别,表明现有模型对胸部疾病长期演变建模能力有限。
原文摘要 · Abstract (English)
Chest radiograph interpretation requires temporal reasoning over prior and current studies, yet most vision-language models are trained on static image-report pairs and lack explicit supervision for modeling longitudinal change. We introduce CheXTemporal, a dataset for temporally grounded reasoning in chest radiography consisting of paired prior-current chest X-rays (CXR) with finding-level temporal and spatial annotations. The dataset includes a five-class progression taxonomy (new, worse, stable, improved, resolved), localized spatial supervision of pathology, explicit spatial-temporal alignment across paired studies, and multi-source coverage for cross-domain evaluation. We additionally construct a 280K-pair silver dataset with automatically derived temporal and anatomical supervision for large-scale evaluation under weaker supervision. Using these resources, we evaluate multiple state-of-the-art vision-language CXR models on grounding and progression-classification tasks in a zero-shot setting. Across both gold and silver evaluations, current models exhibit consistent limitations in spatial grounding, fine-grained temporal reasoning, and robustness under distribution shift. In particular, models perform substantially better on salient progression categories such as worse than on temporally subtle states such as stable and resolved, suggesting limited modeling of longitudinal disease evolution in chest radiography.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。