arXiv:2608.11534cs.CLcs.CV2026-08中稿 · COLM

构建纵向CT影像变化报告基准,让AI学会识别病灶随时间演变。

CT-$Δ$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models

论文配图:CT-$Δ$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models
图 1 · 摘自论文原文
  • 设计新基准CT-ΔBench,支持跨时序扫描的患者级划分防泄露
  • 提出关注变化的评估指标,医生验证报告可靠性与事件提取准确率
  • 对比直接差分与两阶段报告生成,验证端到端方案更优,适合临床决策

在医学影像中,计算机断层扫描(CT)的临床价值不仅在于呈现当前病情,更在于通过序列扫描的纵向比较来判断疾病进展,这一过程对疗效评估、复发检测和持续患者管理至关重要。然而,现有医疗基础模型大多局限于单次扫描理解,缺乏对时序交叉分析的支持。为此,本文研究纵向影像差异报告任务:给定同一患者的两期时序扫描,生成描述其间变化的临床报告。我们构建了专门的基准CT-ΔBench,采用患者级划分以防止信息泄露;为超越表面文本相似性,开发了关注变化的评估指标,并通过独立医生验证评估合成参考文本和事件提取管道的可靠性。还比较了直接成对CT推理与先生成单时点报告再进行文本差分的两阶段方法。最后,提出DeltaMed作为直接成对CT差异报告的基线模型,并在其训练集上进行训练。这些工作为具备时序感知能力的医疗基础模型奠定了基础,使其更贴近真实临床中的纵向推理过程。

原文摘要 · Abstract (English)

In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to determine disease evolution, a process that underpins response assessment, recurrence detection, and ongoing patient management. Yet, despite this central role of temporal comparison in clinical decision-making, existing medical foundation models remain largely confined to single-study understanding, leaving temporally grounded cross-examination insufficiently addressed. To address this gap, we study longitudinal imaging difference reporting, a task in which a model takes two temporally separated scans from the same patient and generates a clinically meaningful report describing interval changes between them. We introduce CT-$Δ$Bench, a dedicated benchmark for this task with patient-level splitting to prevent information leakage. To better evaluate this task beyond surface-level text similarity, we further develop change-aware metrics specifically designed to capture clinically meaningful longitudinal changes, and conduct an independent physician validation to assess the reliability of the synthesized references and event extraction pipeline. We also compare direct paired-CT reasoning with an indirect two-stage pipeline that first generates single-timepoint reports and then performs textual differencing. Finally, we propose DeltaMed, a baseline model for direct paired-CT difference reporting, and train it on the benchmark training set. Together, these contributions lay the groundwork for temporally aware medical foundation models that better reflect real-world longitudinal clinical reasoning.

医学影像纵向分析视觉语言模型差异报告

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。