arXiv:2606.16047cs.CL2026-06中稿 · publication in the…被引 1

用多智能体辩论识别论点关系,只在不确定时讨论,效果优于不训练模型。

From Argument Components to Graphs: A Multi-Agent Debate with Confidence Gating for Argument Relations

论文配图:From Argument Components to Graphs: A Multi-Agent Debate with Confidence Gating for Argument Relations
图 1 · 摘自论文原文
  • 让多个智能体辩论论点对,只在置信度低时才讨论
  • 在英国论文数据集上达成最高中文宏平均F1,比基线高3.2个百分点
  • 结果可读性强,适合需要解释性的论证分析场景

大语言模型在论点挖掘领域应用日益广泛,因其具备强大的通用推理能力。然而,无需训练的标准模型常忽略需联合分析的复杂上下文细节,且自我修正机制易强化初始幻觉。克服这些缺陷通常需昂贵的领域特定监督微调。近期研究显示,通过主张-反对-裁判架构的多智能体辩证优化,可在组件分类任务中有效解决此类问题,为无训练方法指明方向。本文将该框架拓展至论点关系识别与分类(ARIC)任务,将问题重构为对论点对的辩论。此外引入置信度门控机制,仅对置信度低的情况进行辩论,高置信时接受初始预测。在UKP Argument Annotated Essays v2数据集上,选择性辩论达到所有无训练方法中的最高宏观F1,而对所有样本进行辩论则使性能低于某一基线。所有生成式方法均在宏观F1上超越微调的RoBERTa模型,表明攻击类样本的欠表示对监督微调的损害大于对仅推理模型的影响。此外,该框架生成人类可读的辩论记录,提供了单智能体和监督分类器所缺乏的可解释性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly assessed and utilized in the field of Argument Mining (AM), thanks to their strong general reasoning capabilities. However, standard training-free models often miss sophisticated details, specifically in contexts where two parts of the text have to be analyzed together. Furthermore, self-correction mechanisms tend to reinforce initial hallucinations in reasoning. Overcoming these limitations typically requires expensive, domain-specific supervised fine-tuning. Recent work has shown that a multi-agent paradigm can address such weaknesses for the component classification task through dialectical refinement with a Proponent-Opponent-Judge architecture, setting a promising direction for training-free approaches in the field. In this paper, we extend and evaluate this framework on the Argument Relation Identification and Classification (ARIC) task, reformulating it as a debate over component pairs. Besides that, we introduce a confidence gating mechanism that enables debating only on the uncertain cases and accepting the initial prediction when confidence is high. On the UKP Argument Annotated Essays v2 corpus, we demonstrate that the selective debate achieves the highest Macro F1 among all training-free methods, while debate over all samples degrades performance below that of one of the baselines. All generative approaches also outperform fine-tuned RoBERTa models on Macro F1, suggesting that the under-representation of the Attack class was more damaging to supervised fine-tuning than to inference-only models. Additionally, our framework produces human-readable debate transcripts, offering interpretability absent from both single-agent and supervised classifiers.

论点挖掘多智能体可解释性推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。