arXiv:2603.27451cs.CLcs.AI2026-03中稿 · publication in the…被引 3

多智能体辩论提升论点分类准确率,无需领域训练。

Multi-Agent Dialectical Refinement for Enhanced Argument Classification

  • 用正方反方裁判三智能体辩论解析模糊论点结构
  • 在学生作文数据集上达85.7%宏平均F1,超越单模型基线
  • 输出可读辩论记录,结果透明可解释,适合教育评估场景

论点挖掘(AM)是自动化写作评价的基础技术,但传统监督方法依赖昂贵的领域特定微调。尽管大语言模型(LLMs)提供无训练替代方案,却常因结构歧义难以区分主张与前提等相似组件。此外,单智能体自我修正机制易陷入自我说服,强化初始错误。本文提出MAD-ACC(多智能体辩论论点成分分类框架),通过辩证精炼解决分类不确定性。该框架采用正方-反方-裁判模型,让智能体对模糊文本的不同解释进行辩护,揭示单模型遗漏的逻辑细节。在UKP学生作文语料库上的评估显示,MAD-ACC实现85.7%的宏平均F1,显著优于单智能体推理基线,且无需领域特定训练。同时,不同于“黑箱”分类器,其辩证方法生成人类可读的辩论记录,清晰展现决策依据,提供透明可解释的替代方案。

原文摘要 · Abstract (English)

Argument Mining (AM) is a foundational technology for automated writing evaluation, yet traditional supervised approaches rely heavily on expensive, domain-specific fine-tuning. While Large Language Models (LLMs) offer a training-free alternative, they often struggle with structural ambiguity, failing to distinguish between similar components like Claims and Premises. Furthermore, single-agent self-correction mechanisms often suffer from sycophancy, where the model reinforces its own initial errors rather than critically evaluating them. We introduce MAD-ACC (Multi-Agent Debate for Argument Component Classification), a framework that leverages dialectical refinement to resolve classification uncertainty. MAD-ACC utilizes a Proponent-Opponent-Judge model where agents defend conflicting interpretations of ambiguous text, exposing logical nuances that single-agent models miss. Evaluation on the UKP Student Essays corpus demonstrates that MAD-ACC achieves a Macro F1 score of 85.7%, significantly outperforming single-agent reasoning baselines, without requiring domain-specific training. Additionally, unlike "black-box" classifiers, MAD-ACC's dialectical approach offers a transparent and explainable alternative by generating human-readable debate transcripts that explain the reasoning behind decisions.

论点挖掘多智能体可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。