arXiv:2608.29738cs.CL2026-08

LLM生成的说服性对话看似有力,实则逻辑薄弱。

Evaluating the Capabilities of LLMs for Persuasive Dialogue

论文配图:Evaluating the Capabilities of LLMs for Persuasive Dialogue
图 1 · 摘自论文原文
  • 构建多智能体辩论平台,用论证理论判定对话胜负
  • 模型主观说服力强但正式论证能力差,人类仍具竞争力
  • 检索增强与多智能体设置加剧了说辞流畅与逻辑严谨的分离

大型语言模型(LLMs)能生成看似极具说服力的文本,但听起来有说服力是否意味着论点扎实?我们提出了 extsc{Persuasio},一个基于形式论证理论的多智能体对话平台,用于在自由文本辩论中裁定逻辑胜者。通过该系统,我们在英国政治议题上生成了192场人类与LLM之间的辩论,并通过自动化裁定和9,702次众包两两评判(共1,386个标注实例)评估了22位对话者。结果发现,主观说服力与形式论证能力之间存在持续脱节:尽管LLM在主观排名中占优,但在论证理论判定下表现明显较差,而人类依然具备竞争力。多智能体与检索增强版本进一步扩大了这一差异。这些发现揭示了基于LLM的说服性对话中,修辞流畅性与正式论证力之间存在系统性鸿沟。

原文摘要 · Abstract (English)

Large language models (LLMs) can generate apparently highly persuasive text, but does sounding persuasive mean arguing well? We introduce \textsc{Persuasio}, a multi-agent dialogue platform grounded in a formal argumentation-based theory of persuasion dialogues that adjudicates logical winners during free-text debates. Using this system, we generated 192 debates on a UK political topic between humans and LLMs, and evaluated 22 interlocutors through both automated adjudication and 9,702 crowdsourced pairwise judgements across 1{,}386 annotation instances. We observed a consistent decoupling between subjective and formal persuasiveness: LLMs dominated the subjective ranking yet performed substantially worse under argumentation-theoretic adjudication, where humans remained competitive. Multi-agent and retrieval-augmented variants further widened this divergence. These findings reveal a systematic gap between rhetorical fluency and formal argumentative strength in LLM-based persuasive dialogues.

说服对话论证理论LLM评估多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。