arXiv:2510.25110cs.CL2025-10被引 4

构建大规模基准,评估大模型角色扮演在观点演化中的真实度。

DEBATE: A Large-Scale Benchmark for Evaluating Opinion Dynamics in Role-Playing LLM Agents

  • 用真实人类对话与态度数据构建多轮辩论基准
  • 大模型模拟组倾向过早收敛,与真人差异明显
  • 适合研究社会极化、虚假信息的可信仿真

准确建模社交互动中的观点演变对理解并缓解极化、误导信息与社会冲突至关重要。现有基于角色扮演大模型代理(RPLA)的多智能体模拟常出现不自然群体行为,如过早收敛,且缺乏与真实人类互动对齐的实证基准。本文提出DEBATE,一个大规模基准,用于评估多智能体RPLA模拟中观点动态的真实性。该基准包含来自美国107个话题的2,788名参与者、697个群体的多轮公开言论与私密李克特量表态度数据,支持话语级与群体级评估,并为个体级分析预留空间。我们使用七种大模型构建‘数字孪生’代理,在两个场景下评估:下一条消息预测与完整动态模拟,采用基于立场的观点演化指标。零样本条件下,RPLA群体表现出强于真人组的观点收敛。在保留组划分上,对Llama-3.1-8B-Instruct进行监督微调(SFT)可提升辅助立场对齐并降低群体层面收敛误差,但观点变化与信念更新仍存差异。DEBATE为模拟观点动态提供了严谨评测框架,推动未来研究实现多智能体代理与真实人类互动对齐。

原文摘要 · Abstract (English)

Accurately modeling opinion change through social interactions is crucial for understanding and mitigating polarization, misinformation, and societal conflict. Recent work simulates opinion dynamics with role-playing LLM agents (RPLAs), but multi-agent simulations often display unnatural group behavior, such as premature convergence, and lack empirical benchmarks for assessing alignment with real human group interactions. We introduce DEBATE, a large-scale benchmark for evaluating the authenticity of opinion dynamics in multi-agent RPLA simulations. DEBATE contains multi-round public messages and private Likert-scale beliefs from U.S.-based participants across 107 topics; the cleaned benchmark used in our experiments contains 2,788 participants in 697 groups, enabling evaluation at the utterance and group levels and supporting future individual-level analyses. We instantiate "digital twin" RPLAs with seven LLMs and evaluate across two settings: next-message prediction and full dynamics simulation, using stance-based opinion-dynamics metrics. In zero-shot settings, RPLA groups exhibit strong opinion convergence relative to human groups. On the held-out group split, supervised fine-tuning (SFT) for Llama-3.1-8B-Instruct improves auxiliary stance alignment and reduces group-level convergence error, though discrepancies in opinion change and belief updating remain. DEBATE enables rigorous benchmarking of simulated opinion dynamics and supports future research on aligning multi-agent RPLAs with realistic human interactions.

角色扮演观点演化大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。