LLMs在民主讨论中缺乏深度思辨能力,难以替代人类集体决策。
The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse
- 用政治学中的协商质量指数评估12个议题下1980次LLM对话
- LLM群体观点多样性仅为人类的三分之一,且越讨论越分散
- 适合辅助人类决策,但不宜独立承担公共议题协商
大型语言模型(LLMs)正被越来越多地用于需要对复杂、价值导向问题进行集体推理的场景。当前对这些应用的信心主要基于可验证任务(如数学、编程、协调游戏)的基准测试,然而许多实际应用场景并不存在客观正确答案,决策质量取决于整合多元视角以达成共识。我们主张,仅凭可验证任务的基准无法充分推断LLM在这一类问题上的推理能力,且对话语过程的评估(如尊重性、论证合理性、参与度)存在系统性不足。本文采用政治科学中已验证的协商理性指数(Deliberative Reason Index, DRI),评估1,980次五代理论模型在12个公民议事议题上的表现。结果表明,尽管LLM群体的对话程序质量与人类相当,但跨主体一致性提升有限,且仅在易处理而非伦理争议性强的问题上显著;其观点多样性约为人类的三分之一,且与人类相反——人类讨论后趋于收敛,而LLM讨论后分歧加剧。通过角色提示增强多样性并未恢复人类模式,反而改变了协商更新机制。结论是:当前证据不支持将LLM视为自主的协商主体,它们更适合作为辅助工具支持人类推理。
原文摘要 · Abstract (English)
LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer exists and where decision quality instead depends on integrating pluralistic perspectives to find mutually acceptable solutions. We argue that LLM reasoning capacity on this class of problems cannot be fully inferred from verifiable-task benchmarks, and that procedural evaluations of LLM discourse (respectfulness, justification, engagement) are systematically insufficient. We apply the Deliberative Reason Index (DRI), a measure developed in political science and validated across citizen assemblies, as a tool for evaluating reliable group-level reasoning on pluralistic, non-verifiable problems. Synthesizing recent evidence across 1,980 five-agent LLM runs on 12 citizen-assembly topics across 11 frontier model configurations, we find that LLM groups produce discourse with procedural quality comparable to human deliberation, while gains in intersubjective consistency are small, topic-dependent, and concentrated on tractable rather than ethically contested questions. LLM groups exhibit roughly one-third the perspective diversity of human assemblies and reverse the human convergence pattern: human deliberation decreases dispersion as diverse views synthesise, whereas LLM deliberation increases it. Engineering diversity through persona prompting does not restore the human dynamic but inverts which component of deliberative reasoning is updated. Our conclusion is constraining rather than prohibitive: LLMs can function as tools supporting human reasoning on pluralistic problems, but current evidence does not license treating them as autonomous deliberative agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。