arXiv:2605.07646cs.CLcs.AI2026-05被引 1

让大模型像专家辩论一样分步验证,提升推理可信度。

MAVEN: Multi-Agent Verification-Elaboration Network with In-Step Epistemic Auditing

论文配图:MAVEN: Multi-Agent Verification-Elaboration Network with In-Step Epistemic Auditing
图 1 · 摘自论文原文
  • 用角色分离的辩论机制,分步骤验证推理过程。
  • 在多个评测中优于现有模型,关键指标提升显著。
  • 不依赖特定模型,可通用增强各类大模型推理能力。

尽管显式推理轨迹能提升模型可解释性,但现有方法多采用单一链式结构,缺乏中间验证,导致早期错误持续传播。这种非模块化设计阻碍了精细审计,影响高风险应用中的认知信任。本文提出MAVEN(多智能体验证-扩展网络与阶段内认知审计),受黑板系统启发,通过角色解耦将大模型转化为有意识的推理者。核心是模拟专家辩论的对抗性怀疑者-研究者-裁判循环,功能上分离逻辑辩护与事实依据。在OpenBookQA、TruthfulQA、HALUEVAL和StrategyQA四个基准上的实验表明,MAVEN在四项细粒度指标上均表现更优。尤其在生成显式结构化、模块化且可验证的推理路径方面,显著超越基于隐式推理的模型(如GEMINI-3.1-Pro)及后验共识基线(如ReConcile)。全面评估还证实,MAVEN具备完全的模型无关性,可作为强而通用的推理增强器,在多种骨干模型上带来显著性能提升。

原文摘要 · Abstract (English)

While explicit reasoning trajectories enhance model interpretability, existing paradigms often rely on monolithic chains that lack intermediate verification, allowing early errors to cascade unchecked. This lack of modularity impedes granular auditing and compromises the epistemic trust required for high-stakes applications. We propose MAVEN (Multi-Agent Verification-Elaboration Network with In-Step Epistemic Auditing), a blackboard-inspired framework designed to transform LLMs into deliberate reasoners through explicit role-decoupling. At its core, MAVEN operationalizes an adversarial Skeptic-Researcher-Judge loop, simulating expert deliberation by functionally separating logical defense from factual grounding. Experiments on OpenBookQA, TruthfulQA, HALUEVAL and StrategyQA benchmarks demonstrate that MAVEN delivers superior reasoning quality across four fine-grained metrics. Notably, MAVEN consistently outperforms latent reasoning models such as GEMINI-3.1-Pro and consensus-based baselines (e.g., ReConcile) by generating explicitly structured, modular, and verifiable deliberation trajectories, rather than relying on implicit internal states or post-hoc consensus. Moreover, comprehensive evaluations confirm that MAVEN is fully model-agnostic, serving as a strong and transferable reasoning booster that yields substantial performance improvements across diverse backbone models.

推理增强多智能体可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。