arXiv:2601.15129cs.CL2026-01被引 2

构建200张胸部X光片权威标注数据集,助力医学AI模型评估。

RSNA Large Language Model Benchmark Dataset for Chest Radiographs of Cardiothoracic Disease: Radiologist Evaluation and Validation Enhanced by AI Labels (REVEAL-CXR)

  • 用AI辅助筛选并由放射科医生三重验证,确保标注质量。
  • 创建100张公开+100张保留的基准数据集,含12类病灶标签。
  • 适合医学AI研究者用于模型评测与可解释性分析。

多模态大语言模型在选择题式影像科考试中已达到住院医师水平。但要开发临床可用的多模态LLM工具,需依赖领域专家构建高质量基准数据集。本研究基于MIDRC数据库的13,735张去标识化胸部X光片及报告,利用GPT-4o提取异常发现,并通过本地部署的Phi-4-Reasoning模型映射至12个基准标签。从中随机抽取1,000例,依据AI建议标签进行临床相关性与难度分布均衡采样,供17名胸片放射科医生评审。每位图像由三位专家评估,判断是否“完全同意”、“基本同意”或“不同意”模型建议标签。最终,381张获得至少两名医生“完全同意”的图像被选中,其中200张进一步筛选,优先包含罕见或多病灶情况,分为100张公开数据和100张保留测试集(仅由RSNA用于独立模型评估)。该200例基准数据集已公开发布于https://imaging.rsna.org,每例均经三名医生验证。同时提出一种AI辅助标注流程,提升标注效率,减少遗漏,支持半协同工作环境。

原文摘要 · Abstract (English)

Multimodal large language models have demonstrated comparable performance to that of radiology trainees on multiple-choice board-style exams. However, to develop clinically useful multimodal LLM tools, high-quality benchmarks curated by domain experts are essential. To curate released and holdout datasets of 100 chest radiographic studies each and propose an artificial intelligence (AI)-assisted expert labeling procedure to allow radiologists to label studies more efficiently. A total of 13,735 deidentified chest radiographs and their corresponding reports from the MIDRC were used. GPT-4o extracted abnormal findings from the reports, which were then mapped to 12 benchmark labels with a locally hosted LLM (Phi-4-Reasoning). From these studies, 1,000 were sampled on the basis of the AI-suggested benchmark labels for expert review; the sampling algorithm ensured that the selected studies were clinically relevant and captured a range of difficulty levels. Seventeen chest radiologists participated, and they marked "Agree all", "Agree mostly" or "Disagree" to indicate their assessment of the correctness of the LLM suggested labels. Each chest radiograph was evaluated by three experts. Of these, at least two radiologists selected "Agree All" for 381 radiographs. From this set, 200 were selected, prioritizing those with less common or multiple finding labels, and divided into 100 released radiographs and 100 reserved as the holdout dataset. The holdout dataset is used exclusively by RSNA to independently evaluate different models. A benchmark of 200 chest radiographic studies with 12 benchmark labels was created and made publicly available https://imaging.rsna.org, with each chest radiograph verified by three radiologists. In addition, an AI-assisted labeling procedure was developed to help radiologists label at scale, minimize unnecessary omissions, and support a semicollaborative environment.

医学影像大模型评测胸部X光专家标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。