构建大规模胸片报告纠错数据集,评估并提升AI报告修正能力。
CorBenchX: Large-Scale Chest X-Ray Error Dataset and Vision-Language Model Benchmark for Report Error Correction
- 用大模型生成2.6万份带常见错误的胸片报告,配原始文本与错误标注。
- o4-mini在零样本下纠错效果最佳,但整体仍未达临床可用水平。
- 提出多步强化学习框架,使开源模型纠错精度提升超38%。
AI模型在检测放射科报告错误方面展现出巨大潜力,但缺乏统一的评估基准。为此,我们推出了CorBenchX,一个用于胸部X光报告自动错误检测与修正的综合性评测体系。我们通过提示DeepSeek-R1注入临床常见错误,合成包含26,326份带错报告的数据集,每份均配有原始文本、错误类型及人工可读说明。基于此数据集,我们对InternVL、Qwen-VL、GPT-4o、o4-mini和Claude-3.7等开闭源视觉语言模型进行零样本提示下的错误检测与修正评测。其中,o4-mini表现最优,检测准确率达50.6%,纠正评分包括BLEU 0.853、ROUGE 0.924、BERTScore 0.981、SembScore 0.865和CheXbertF1 0.954,但仍低于临床应用标准,凸显精准修正的挑战。为推动技术进步,我们提出多步强化学习(MSRL)框架,融合格式合规性、错误类型准确率与BLEU相似度的多目标奖励。将MSRL应用于基准中表现最佳的开源模型QwenVL2.5-7B,使单错误检测精度提升38.3%,单错误修正性能提升5.2%。
原文摘要 · Abstract (English)
AI-driven models have shown great promise in detecting errors in radiology reports, yet the field lacks a unified benchmark for rigorous evaluation of error detection and further correction. To address this gap, we introduce CorBenchX, a comprehensive suite for automated error detection and correction in chest X-ray reports, designed to advance AI-assisted quality control in clinical practice. We first synthesize a large-scale dataset of 26,326 chest X-ray error reports by injecting clinically common errors via prompting DeepSeek-R1, with each corrupted report paired with its original text, error type, and human-readable description. Leveraging this dataset, we benchmark both open- and closed-source vision-language models,(e.g., InternVL, Qwen-VL, GPT-4o, o4-mini, and Claude-3.7) for error detection and correction under zero-shot prompting. Among these models, o4-mini achieves the best performance, with 50.6 % detection accuracy and correction scores of BLEU 0.853, ROUGE 0.924, BERTScore 0.981, SembScore 0.865, and CheXbertF1 0.954, remaining below clinical-level accuracy, highlighting the challenge of precise report correction. To advance the state of the art, we propose a multi-step reinforcement learning (MSRL) framework that optimizes a multi-objective reward combining format compliance, error-type accuracy, and BLEU similarity. We apply MSRL to QwenVL2.5-7B, the top open-source model in our benchmark, achieving an improvement of 38.3% in single-error detection precision and 5.2% in single-error correction over the zero-shot baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。