arXiv:2506.19217cs.CVcs.AI2025-06被引 2

构建首个用于检测修正CT报告错误的医学视觉问答基准

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports

  • 设计六类错误分类的多层级视觉问答任务
  • 发现主流3D医疗大模型在不同错误类型上表现差异显著
  • 适合医学AI验证、临床辅助诊断系统研发者使用

计算机断层扫描(CT)在临床诊断中至关重要,但检查量增加也带来了诊断错误风险。尽管多模态大语言模型在医学知识理解上表现出色,却常产生不准确信息,亟需严格验证。现有医学视觉问答基准多聚焦简单视觉识别,缺乏临床相关性且无法评估专家级知识。本文提出MedErr-CT,一个基于VQA框架、用于评估模型识别与修正CT报告错误能力的新基准。该基准包含六类错误:四类视觉相关错误(遗漏、误增、方向、尺寸)和两类词汇错误(单位、拼写),分为分类、检测、修正三阶任务。通过该基准,我们定量评估了前沿3D医疗大模型性能,发现其在不同错误类型间表现差异显著。本基准有助于推动更可靠、可临床应用的医疗大模型发展,最终减少诊断误差,提升诊疗准确性。代码与数据集见https://github.com/babbu3682/MedErr-CT。

原文摘要 · Abstract (English)

Computed Tomography (CT) plays a crucial role in clinical diagnosis, but the growing demand for CT examinations has raised concerns about diagnostic errors. While Multimodal Large Language Models (MLLMs) demonstrate promising comprehension of medical knowledge, their tendency to produce inaccurate information highlights the need for rigorous validation. However, existing medical visual question answering (VQA) benchmarks primarily focus on simple visual recognition tasks, lacking clinical relevance and failing to assess expert-level knowledge. We introduce MedErr-CT, a novel benchmark for evaluating medical MLLMs' ability to identify and correct errors in CT reports through a VQA framework. The benchmark includes six error categories - four vision-centric errors (Omission, Insertion, Direction, Size) and two lexical error types (Unit, Typo) - and is organized into three task levels: classification, detection, and correction. Using this benchmark, we quantitatively assess the performance of state-of-the-art 3D medical MLLMs, revealing substantial variation in their capabilities across different error types. Our benchmark contributes to the development of more reliable and clinically applicable MLLMs, ultimately helping reduce diagnostic errors and improve accuracy in clinical practice. The code and datasets are available at https://github.com/babbu3682/MedErr-CT.

医学AI视觉问答错误检测医疗大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。