arXiv:2507.01936cs.CLcs.CY2025-07ACL被引 6

LLMs能说服人却不懂深层论点,对话能力与理解力脱节。

The Thin Line Between Comprehension and Persuasion in LLMs

  • 通过人类与LLM辩论对比,评估其说服力与理解力
  • LLM能维持连贯辩论并影响双方信念,但无法识别论点质量
  • 适合关注AI伦理、对话系统可信度的研究者阅读

大型语言模型(LLMs)在维持高水准、有说服力的对话方面表现优异,但其说服力是否反映对话语的真实理解仍不明确。我们通过人类与LLM之间的非正式辩论来检验这一问题,首先衡量其说服能力,再将其与对论点结构和语用背景的理解能力相关联。结果发现,LLM能有效维持连贯且具说服力的辩论,并影响参与者及观众的信念;但当意识到AI参与时,人们会更批判性地看待其论证。然而,我们还发现LLM无法理解深层对话结构,如论点质量或前提支持关系。研究揭示了LLM理解能力与对话技能之间的断层,对解释关键场景中的部署带来伦理与实践挑战。从论证理论视角出发,实验质疑:一个能令人信服地展开对话的代理,是否必须真正理解所讨论的内容。

原文摘要 · Abstract (English)

Large language models (LLMs) are excellent at maintaining high-level, convincing dialogue, but it remains unclear whether their persuasive success reflects genuine understanding of the discourse. We examine this question through informal debates between humans and LLMs, first by measuring their persuasive skills, and then by relating these to their understanding of _what_ is being talked about: namely, their comprehension of argumentative structures and the pragmatic context on the same debates. We find that LLMs effectively maintain coherent, persuasive debates, and can sway the beliefs of both participants and audiences. We also note that awareness or suspicion of AI involvement encourage people to be more critical of the arguments made. However, we also find that LLMs are unable to show comprehension of deeper dialogical structures, such as argument quality or existence of supporting premises. Our results reveal a disconnect between LLM comprehension and dialogical skills, raising ethical and practical concerns on their deployment on explanation-critical contexts. From an argumentation-theoretical perspective, we experimentally question whether an agent, if it can convincingly maintain a dialogue, is required to show it knows what is talking about.

大模型推理对话理解说服力分析人工智能伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。