构建胸部X光片病变感知基准,评估模型像专家一样看图的能力。
CheXpercept: A Benchmark for Evaluating Expert-Level Lesion Perception in Chest X-rays

- 按放射科医生思维流程设计三阶感知任务:粗略检测、轮廓判断、属性提取。
- 14个模型在细粒度任务上准确率大幅下降,医学专用模型无明显优势。
- 6名专家审核,含2100张片子的10400个问答,覆盖7类关键病变。
当前胸部X光影像的视觉-语言模型评估多局限于疾病存在性分类,缺乏视觉定位能力验证,难以确保临床可靠性。为此,我们提出CheXpercept,一个模拟放射科医生认知流程的多层次感知基准,涵盖粗粒度检测、细粒度轮廓评估与修正、以及语义属性提取三个阶段。为保证大规模下临床真实性,数据通过半自动化生成流程并经六位医学专家审核。该数据集包含2100张胸部X光片,生成10400个问答对,覆盖七种临床关键肺部和心脏病变。我们在该基准上评测14个通用及医学领域视觉-语言模型,发现模型仅在粗粒度层面表现尚可,深入视觉任务时准确率急剧下降;值得注意的是,医学专用模型与通用模型相比几乎无感知优势,暴露出当前领域适配的系统性缺陷。代码与数据集将公开。
原文摘要 · Abstract (English)
The evaluation of vision-language models (VLMs) for chest X-ray (CXR) analysis has largely been limited to disease-presence classification without visual grounding. Such evaluations fail to verify the expert-level lesion perception necessary to ensure the clinical reliability of VLMs. To address these limitations, we introduce CheXpercept, a sequential, multi-level perception benchmark that mirrors a radiologist's cognitive workflow across coarse-level detection, fine-level contour evaluation and revision, and semantic-level attribute extraction. To ensure high clinical fidelity at scale, we construct the dataset using a semi-automated generation pipeline paired with a review by six medical experts. CheXpercept contains 10,400 QA items derived from 2,100 CXRs, covering seven clinically critical pulmonary and cardiac lesions. To demonstrate the current landscape of VLM perception, we benchmark 14 general and medical VLMs on CheXpercept. The models achieve adequate performance only at the coarse level, with accuracy degrading precipitously on deeper visual tasks. Notably, medical VLMs show almost no perceptual advantage over their general-domain counterparts, highlighting a systemic flaw in current domain adaptation. The code and dataset will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。