arXiv:2601.16685cs.AI2026-01

用多智能体模拟放射科医生诊断流程,让报告生成评估更可信。

AgentsEval: Clinically Faithful Evaluation of Medical Imaging Reports via Multi-Agent Reasoning

  • 分步骤模拟医生诊断:定义标准、提取证据、对齐信息、打分
  • 在5个数据集上验证,对改写和语义变化仍保持稳定可靠
  • 适合需要可解释医疗评估的AI研发与临床部署团队

自动医学影像报告的临床正确性与推理一致性评估仍是关键挑战。现有方法难以捕捉放射科诊断中的结构化逻辑,导致判断不可靠且缺乏临床意义。我们提出AgentsEval,一种多智能体流式推理框架,模拟放射科医生的协作诊断流程。将评估过程分解为可解释的步骤:标准定义、证据提取、对齐分析与一致性评分,提供清晰推理路径与结构化临床反馈。同时构建涵盖五个医学报告数据集的多领域扰动基准,覆盖多种成像模态并引入可控语义变化。实验表明,AgentsEval能实现与临床一致、语义忠实且可解释的评估,在同义改写、语义扰动和风格变化下仍保持稳健。该框架推动了透明化、临床导向的报告生成系统评估,助力大语言模型在临床实践中的可信集成。

原文摘要 · Abstract (English)

Evaluating the clinical correctness and reasoning fidelity of automatically generated medical imaging reports remains a critical yet unresolved challenge. Existing evaluation methods often fail to capture the structured diagnostic logic that underlies radiological interpretation, resulting in unreliable judgments and limited clinical relevance. We introduce AgentsEval, a multi-agent stream reasoning framework that emulates the collaborative diagnostic workflow of radiologists. By dividing the evaluation process into interpretable steps including criteria definition, evidence extraction, alignment, and consistency scoring, AgentsEval provides explicit reasoning traces and structured clinical feedback. We also construct a multi-domain perturbation-based benchmark covering five medical report datasets with diverse imaging modalities and controlled semantic variations. Experimental results demonstrate that AgentsEval delivers clinically aligned, semantically faithful, and interpretable evaluations that remain robust under paraphrastic, semantic, and stylistic perturbations. This framework represents a step toward transparent and clinically grounded assessment of medical report generation systems, fostering trustworthy integration of large language models into clinical practice.

医疗报告多智能体可解释性临床评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。