arXiv:2602.14879cs.CVcs.AI2026-02被引 2

首个面向CT病变理解的多模态基准,助力AI精准识别与描述病灶。

CT-Bench: A Benchmark for Multimodal Lesion Understanding in Computed Tomography

  • 构建包含2万+病灶的标注数据集,含边界框、描述与尺寸信息。
  • 涵盖2850个问答对,覆盖定位、描述、大小估计等多任务评估。
  • 支持模型微调,显著提升性能,适合医学AI研究者使用。

人工智能可自动勾画CT图像中的病灶并生成放射科报告内容,但受限于缺乏公开的带病灶级标注的CT数据集。为填补这一空白,我们提出CT-Bench,首个综合性基准数据集,包含两部分:1)病灶图像与元数据集,涵盖7,795例CT检查中的20,335个病灶,提供边界框、描述和尺寸信息;2)多任务视觉问答基准,含2,850个问答对,覆盖病灶定位、描述、尺寸估计和属性分类,其中包含困难负样本以模拟真实诊断挑战。我们评估了多种先进多模态模型(包括视觉-语言模型与医学CLIP变体),并与放射科医生评估结果对比,验证了CT-Bench作为病灶分析综合基准的价值。此外,在病灶图像与元数据集上微调模型,在两个组件上均取得显著性能提升,凸显其临床应用潜力。

原文摘要 · Abstract (English)

Artificial intelligence (AI) can automatically delineate lesions on computed tomography (CT) and generate radiology report content, yet progress is limited by the scarcity of publicly available CT datasets with lesion-level annotations. To bridge this gap, we introduce CT-Bench, a first-of-its-kind benchmark dataset comprising two components: a Lesion Image and Metadata Set containing 20,335 lesions from 7,795 CT studies with bounding boxes, descriptions, and size information, and a multitask visual question answering benchmark with 2,850 QA pairs covering lesion localization, description, size estimation, and attribute categorization. Hard negative examples are included to reflect real-world diagnostic challenges. We evaluate multiple state-of-the-art multimodal models, including vision-language and medical CLIP variants, by comparing their performance to radiologist assessments, demonstrating the value of CT-Bench as a comprehensive benchmark for lesion analysis. Moreover, fine-tuning models on the Lesion Image and Metadata Set yields significant performance gains across both components, underscoring the clinical utility of CT-Bench.

医学影像多模态病灶分析基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。