arXiv:2512.24684cs.CLcs.AI2025-12中稿 · eed by AAMAS 2026 …被引 4

用检索增强记忆生成连贯多轮辩论,让AI论点更一致有证据。

R-Debater: Retrieval-Augmented Debate Generation through Argumentative Memory

  • 基于论据记忆构建辩论框架,通过检索历史论点与证据
  • 在7个领域32组辩论中,多轮表现优于主流大模型
  • 适合需要逻辑严谨、立场一致的智能辩论系统研究者

我们提出R-Debater,一种基于论据记忆的多轮辩论生成智能体框架。该系统融合辩论知识库以检索案例证据和历史辩论动作,并通过角色化智能体生成跨轮次连贯话语。在标准化ORCHID辩论数据集上评估,构建了包含1,000项的检索语料库及32组跨7个领域的保留测试集。评估任务包括下一话语生成(以InspireScore衡量:主观性、逻辑性、事实性)和对抗性多轮模拟(以Debatrix评分:论点、来源、语言、总体质量)。相比强基线大模型,R-Debater在单轮与多轮表现均更优。20位经验辩手的人工评估进一步验证其立场一致性与证据使用能力,表明检索增强与结构化规划结合可生成更忠实、立场一致且连贯的多轮辩论。

原文摘要 · Abstract (English)

We present R-Debater, an agentic framework for generating multi-turn debates built on argumentative memory. Grounded in rhetoric and memory studies, the system views debate as a process of recalling and adapting prior arguments to maintain stance consistency, respond to opponents, and support claims with evidence. Specifically, R-Debater integrates a debate knowledge base for retrieving case-like evidence and prior debate moves with a role-based agent that composes coherent utterances across turns. We evaluate on standardized ORCHID debates, constructing a 1,000-item retrieval corpus and a held-out set of 32 debates across seven domains. Two tasks are evaluated: next-utterance generation, assessed by InspireScore (subjective, logical, and factual), and adversarial multi-turn simulations, judged by Debatrix (argument, source, language, and overall). Compared with strong LLM baselines, R-Debater achieves higher single-turn and multi-turn scores. Human evaluation with 20 experienced debaters further confirms its consistency and evidence use, showing that combining retrieval grounding with structured planning yields more faithful, stance-aligned, and coherent debates across turns.

辩论生成检索增强智能代理论据记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。