arXiv:2508.02808cs.CLcs.AI2025-08被引 2

用智能体对话评估放射科报告,让生成结果更可信可解释。

Clinically Grounded Agent-based Report Evaluation: An Interpretable Metric for Radiology Report Generation

  • 通过两个智能体互问互答,基于临床问题判断报告质量。
  • 评估结果与专家判断高度一致,优于传统指标。
  • 适合医学AI开发者和临床研究人员使用。

放射影像在诊断、治疗规划和临床决策中至关重要。视觉-语言基础模型推动了自动化放射科报告生成(RRG)的发展,但安全部署需要可靠的临床评估。现有评估指标常依赖表面相似性或为黑箱,缺乏可解释性。我们提出ICARE(可解释且临床基准的智能体报告评估),利用大语言模型智能体和动态多选题问答(MCQA)实现可解释评估。两个智能体分别基于真实报告或生成报告,提出具有临床意义的问题并互相问答。答案的一致性反映发现的保留与一致性,作为临床精确率与召回率的可解释代理。通过将评分关联到具体问答对,ICARE实现透明、可解释的评估。临床专家研究显示,ICARE与专家判断显著更一致。扰动分析证实其对临床内容敏感且可复现,模型比较揭示了可解释的错误模式。

原文摘要 · Abstract (English)

Radiological imaging is central to diagnosis, treatment planning, and clinical decision-making. Vision-language foundation models have spurred interest in automated radiology report generation (RRG), but safe deployment requires reliable clinical evaluation of generated reports. Existing metrics often rely on surface-level similarity or behave as black boxes, lacking interpretability. We introduce ICARE (Interpretable and Clinically-grounded Agent-based Report Evaluation), an interpretable evaluation framework leveraging large language model agents and dynamic multiple-choice question answering (MCQA). Two agents, each with either the ground-truth or generated report, generate clinically meaningful questions and quiz each other. Agreement on answers captures preservation and consistency of findings, serving as interpretable proxies for clinical precision and recall. By linking scores to question-answer pairs, ICARE enables transparent, and interpretable assessment. Clinician studies show ICARE aligns significantly more with expert judgment than prior metrics. Perturbation analyses confirm sensitivity to clinical content and reproducibility, while model comparisons reveal interpretable error patterns.

医学AI可解释性报告生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。