用分层智能体辩论优化不确定领域的信念,让AI推理更透明可信。
CHAL: Council of Hierarchical Agentic Language

- 构建分层智能体辩论框架,以信念图结构实现可解释的推理更新。
- 实验证明价值体系决定辩论轨迹,多样性提升全体认知精度。
- 适合追求可审计、对齐人类价值观的AI系统研发者。
多智能体辩论已成为提升大模型在确定性任务上推理能力的有前景方法,但现有方法存在结构缺陷:辩论易引发信念路径的鞅过程,多数收益来自多数投票,且大模型在多轮中表现出信心膨胀而非校准。我们认为,辩论及辩证系统的真实价值不在于确定性任务,而在于可证伪领域——任何立场理论上都可能被更强推理推翻。本文提出分层智能体语言理事会(CHAL),一个将可证伪论证作为信念优化引擎的多智能体辩证框架。每个智能体维护一个具有贝叶斯启发架构的CHAL信念图谱(CBS),通过信念论点强度作为可微目标,实现梯度引导的信念修正。元认知价值体系涵盖认识论、逻辑与伦理,并作为可配置超参数控制智能体推理与裁决结果。一系列消融实验表明:裁决者的价值体系决定信念空间中的整体演化轨迹,理事会多样性提升所有参与者的信念质量,框架具备跨领域泛化能力。据我们所知,CHAL是首个将多智能体辩论视为可证伪领域中结构化信念优化的框架。其可审计的信念产物为专门评估可证伪论证奠定了基础,对构建可透明、对齐且受人类监督的AI系统具有深远意义。
原文摘要 · Abstract (English)
Multi-agent debate has emerged as a promising approach for improving LLM reasoning on ground-truth tasks, yet current methodologies face certain structural limitations: debate tends to induce a martingale over belief trajectories, majority voting accounts for most observed gains, and LLMs exhibit confidence escalation rather than calibration across rounds. We argue that the genuine value of debate, and dialectic systems as a whole, lies not in ground-truth tasks but in defeasible domains, where every position can in principle be defeated by better reasoning. We present the Council of Hierarchical Agentic Language (CHAL), a multi-agent dialectic framework that treats defeasible argumentation as an engine for belief optimization. Each agent maintains a CHAL Belief Schema (CBS), a graph-structured belief representation with a Bayesian-inspired architecture, that facilitates belief revision through a gradient-informed dynamic mechanism by leveraging the strength of the belief's thesis as a differentiable objective. Meta-cognitive value systems spanning epistemology, logic, and ethics are elevated to configurable hyperparameters governing agent reasoning and adjudication outcomes. We provide a series of ablation experiments that demonstrate systematic and interpretable effects: the adjudicator's value system determines the debate's overall trajectories in latent belief space, council diversity refines beliefs for all participants, and the framework generalizes across broad fields. CHAL is, to our knowledge, the first framework to treat multi-agent debate as structured belief optimization over defeasible domains. Further, the auditable belief artifacts it produces establish the foundation for dedicated evaluation suites for defeasible argumentation, with broader implications for building AI systems whose reasoning and value commitments are transparent, aligned, and subject to human oversight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。