arXiv:2603.28488cs.CLcs.AI2026-03被引 4

用法庭辩论式框架提升大模型对争议声明的验证准确率

Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification

  • 设计角色分工与渐进式检索,动态扩展证据池
  • 零样本测试达81.7%准确率,比普通多智能体高10个百分点
  • 适合需要可信验证的AI系统,如事实核查与决策支持

大语言模型在高风险声明验证中仍因幻觉和浅层推理不可靠。尽管检索增强生成(RAG)和多智能体辩论(MAD)有所改善,但受限于单轮检索和无序辩论。我们提出法庭式多智能体框架PROClaim,将验证重构为结构化、对抗性讨论。通过引入原告、辩护方、法官等专责角色,结合渐进式RAG(P-RAG),在辩论过程中动态拓展并精炼证据库。同时采用证据协商、自我反思及异构多法官聚合,实现校准、鲁棒性和多样性。在Check-COVID基准的零样本评估中,PROClaim达到81.7%准确率,相较标准多智能体辩论提升10.0个百分点,其中P-RAG贡献主要增益(+7.5个百分点)。结果表明,结构化审议与模型异质性可有效缓解系统性偏差,为可靠声明验证提供坚实基础。代码与数据已公开于https://github.com/mnc13/PROClaim。

原文摘要 · Abstract (English)

Large language models (LLMs) remain unreliable for high-stakes claim verification due to hallucinations and shallow reasoning. While retrieval-augmented generation (RAG) and multi-agent debate (MAD) address this, they are limited by one-pass retrieval and unstructured debate dynamics. We propose a courtroom-style multi-agent framework, PROClaim, that reformulates verification as a structured, adversarial deliberation. Our approach integrates specialized roles (e.g., Plaintiff, Defense, Judge) with Progressive RAG (P-RAG) to dynamically expand and refine the evidence pool during the debate. Furthermore, we employ evidence negotiation, self-reflection, and heterogeneous multi-judge aggregation to enforce calibration, robustness, and diversity. In zero-shot evaluations on the Check-COVID benchmark, PROClaim achieves 81.7% accuracy, outperforming standard multi-agent debate by 10.0 percentage points, with P-RAG driving the primary performance gains (+7.5 pp). We ultimately demonstrate that structural deliberation and model heterogeneity effectively mitigate systematic biases, providing a robust foundation for reliable claim verification. Our code and data are publicly available at https://github.com/mnc13/PROClaim.

大模型验证多智能体信息检索推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。