arXiv:2606.15037cs.CLcs.CV2026-06

用问答方式评估放射科报告,更贴近临床需求。

ReportQA: QA-Based Radiology Report Evaluation

论文配图:ReportQA: QA-Based Radiology Report Evaluation
图 1 · 摘自论文原文
  • 基于临床知识树构建问答对,量化报告生成质量。
  • 新指标QAScore与医生判断更一致,优于传统方法。
  • 适合评估报告生成模型,尤其关注细粒度临床信息。

放射科报告评估对推动自动化报告生成至关重要。现有自然语言生成指标临床相关性弱,临床疗效(CE)指标虽关注重要医学发现,但仅覆盖有限实体且依赖人工标注,难以扩展。临床实践中,报告是信息传递媒介,医生据此完成诊断而无需直接查看图像。为此,我们提出ReportQA框架,支持对报告生成系统的细粒度定量分析。首先收集多模态、多解剖区域数据集;在放射科医生指导下构建临床实体与属性的知识树,并利用大语言模型(LLMs)从原始报告中提取结构化信息;随后通过预定义模板生成问答对,经自过滤和报告过滤进行质量控制。评估时,将报告作为上下文,使用LLM作为裁判模型回答问答对。基于问答准确率,引入QAScore指标。实验表明,该指标与放射科医生判断高度一致。在多个先进视觉-语言模型上测试发现,当前基于报告的推理范式难以学习细粒度临床表征,且存在强烈负向先验偏差;相比之下,问题驱动推理更为有效。为保证可复现性与可扩展性,我们公开知识树、结构化报告、问答对及完整构建与评估代码。

原文摘要 · Abstract (English)

Radiology report evaluation is essential for advancing automated report generation. Natural language generation metrics have limited clinical relevance. Clinical efficacy (CE) metrics evaluate important medical findings, but focus mainly on presence and cover only a limited set of entities. Due to heavy reliance on manual annotations, it is difficult for CE metrics to extend clinical entities or attributes. In clinical practice, radiology reports serve as a medium for information transfer. Clinicians use them to perform downstream diagnostic tasks without directly inspecting images. Based on this insight, we propose ReportQA, a clinical-related and flexible radiology report evaluation framework, supporting detailed quantitative analysis of radiology report generation systems. We first collect datasets covering multiple imaging modalities and anatomical regions. We then construct knowledge trees of clinical entities and attributes with radiologist guidance, and use large language models (LLMs) to extract structured information from raw reports. Next, we generate QA pairs from predefined templates and apply quality control through self-filtering and report-based filtering. During evaluation, the report is treated as context, and an LLM acts as a judge model to answer the QA pairs. Based on the resulting QA accuracy, we introduce QAScore metric. Compared with existing metrics, QAScore shows better alignment with radiologist judgments. Experiments on multiple state-of-the-art vision-language models reveal that current report-based inference paradigms struggle to learn fine-grained clinical representations and exhibit strong negative prior biases. In contrast, question-driven inference provides a more effective alternative. For reproducibility and extensibility, we release the knowledge trees, structured reports, and QA pairs, along with the pipeline code for QA construction and evaluation.

医学影像报告生成评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。