提出后置辩论评判的理论框架,提升AI辩论的可信与可解释性。
A Theory of Post-hoc Debate Judgement

- 构建辩论评判的四大核心属性:可复现、鲁棒、有根据、可解释。
- 对比大模型判官与形式化论证语义,发现二者准确率相近但后者有理论保障。
- 推荐用形式化论证作为可信辩论裁判,适合追求严谨性的AI系统设计者。
辩论已成为增强智能体性能、提升可解释性与用户参与度的有效方法。例如,由大语言模型驱动的智能体可进行内部或外部辩论。在多数应用场景中,辩论结果由外部评判者(常为大模型)事后决定。本文发展并验证了一种适用于所有存在正反论证辩论场景的辩论评判理论。我们识别出评判需满足的若干形式化属性,涵盖可复现性、鲁棒性、根基性和可解释性。通过形式化分析与实验评估,我们在论点验证任务中对比两种评判方法:基于大模型的判官方案和源自计算论证的形式语义方案。结果显示二者准确率相当,但后者具备更强的理论保证。总体表明,论证语义是辩论驱动AI中理想且严谨的评判机制。
原文摘要 · Abstract (English)
Debates have recently emerged as a useful methodology for agentic AI to improve performance as well as to aid explainability and user engagement. For example, LLM-empowered agents may debate internally (with themselves) and/or externally (with other agents). In many settings where debates are used, debates' outcomes and resulting outputs are determined post-hoc by external judges, often LLMs. In this paper we develop and test a novel theory of debate judgement applicable to all settings where agents engage in debates by providing pros and cons for their opinions therein. Specifically, we identify a number of formal properties that debate judgement may be required to satisfy in general, as concerns reproducibility, robustness, groundedness and explainability. Then, we explore their satisfaction formally and/or experimentally, for claim verification settings, for two specific alternative debate judgement methods: variants of the LLMs as a judge idea and formal semantics drawn from computational argumentation. We show that the two methods give similar accuracy performances but the former may lack formal guarantees that the latter brings. Overall, our study indicates argumentation semantics as an ideal candidate for principled judges in debate-driven AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。