arXiv:2606.15419cs.CLcs.AI2026-06中稿 · the Journal of the…被引 1

让大模型互相评题,提升医疗问答的准确性和可信度。

Let LLMs Judge Each Other: Multi-Agent Peer-Reviewed Reasoning for Medical Question Answering

论文配图:Let LLMs Judge Each Other: Multi-Agent Peer-Reviewed Reasoning for Medical Question Answering
图 1 · 摘自论文原文
  • 多个大模型独立推理并互评思路,选最优解
  • 平均准确率82.0%,超越单模型和投票方法
  • 适合追求高可靠性的医学AI应用

目标:提升大语言模型在医疗问答中的准确性、可解释性与鲁棒性。方法:设计多智能体同行评审推理机制,多个大模型独立生成思维链与候选答案,再作为同行评审者评估彼此推理的事实正确性与逻辑合理性,最终选取评分最高的思维链生成答案。在五个先进大模型(Llama-3.1-8B、Qwen2.5-7B、Phi-4、DeepSeek-LLM-7B、GPT-oss-20B)上,于HeadQA、MedQA-USMLE和PubMedQA三个基准数据集上进行实验,对比单模型思维链推理与基于思维链的多数投票方法。结果:同行评审推理持续优于两个基线,最佳模型组合在各数据集上平均准确率达0.820,高于最强单模型(0.777)和多数投票集成(最高0.789)。该方法随参与模型增加而有效扩展,且同行评估能可靠区分高质量与低质量推理链。结论:所提多智能体同行评审推理方法使大模型兼具求解者与评估者角色,在医疗问答中实现更优性能。通过强调推理质量而非仅答案一致性,该方法提升准确性、可解释性与鲁棒性,为可信生物医学AI系统提供新方向。

原文摘要 · Abstract (English)

Objective: To enhance the accuracy, interpretability, and robustness of large language models (LLMs) in medical question answering (MedQA). Method: We designed a multi-agent peer-reviewed reasoning method in which multiple LLM agents independently generate chain-of-thought reasoning with candidate answers, then act as peer reviewers to evaluate each other's reasoning for factual correctness and logical soundness. The highest-rated reasoning chain is selected to produce the final answer. Experiments were conducted with five state-of-the-art LLMs (Llama-3.1-8B, Qwen2.5-7B, Phi-4, DeepSeek-LLM-7B, GPT-oss-20B) on three benchmark datasets: HeadQA, MedQA-USMLE, and PubMedQA. Performance was compared against single-model chain-of-thought reasoning and chain-of-thought-based majority voting. Results: Peer-reviewed reasoning consistently outperformed both baselines. The best model combination achieved an average accuracy of 0.820 across datasets, exceeding the strongest single model (0.777) and majority voting ensembles (up to 0.789). The method also scaled effectively with more participating models, while peer assessments reliably distinguished high- from low-quality reasoning chains. Conclusion: The proposed multi-agent peer-reviewed reasoning method enables LLMs to act as both solvers and evaluators, yielding superior performance in MedQA. By emphasizing reasoning quality rather than answer agreement alone, this approach improves accuracy, interpretability, and robustness, offering a promising direction for trustworthy biomedical AI systems.

医疗问答多智能体推理评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。