arXiv:2506.21587cs.CL2025-06被引 1

对比中美的大模型在民意模拟中的表现,发现深度求索效果更好但仍有偏见。

A Cross-Cultural Comparison of LLM-based Public Opinion Simulation: Evaluating Chinese and U.S. Models on Diverse Societies

  • 用中美真实调查数据测试多个大模型的民意模拟能力。
  • 深求索-V3在美模拟堕胎议题最准,中国模拟外援和个体主义较好。
  • 所有模型都存在群体刻板化问题,需改进训练以减少偏见。

本研究评估了开源大模型DeepSeek在模拟中美公众意见方面的表现,与通义千问2.5、GPT-4o、Llama-3.3等主流模型进行对比。基于美国国家选举研究(ANES)和中国佐标数据集,分析模型对堕胎、气候变化、枪支管控、移民、同性伴侣服务、外国援助、个人主义、资本主义、传统主义及自由市场等议题的意见预测能力。结果显示,DeepSeek-V3在模拟美国堕胎议题上表现最优,主要因其能更准确响应民主党或自由派身份设定;在中国样本中,其在外国援助与个体主义议题上表现最佳,但在资本主义思想上未能捕捉低收入及非大学教育群体立场。所有模型在传统主义与自由市场议题上无显著差异。进一步分析表明,各模型普遍存在对群体内部观点过度同质化的倾向,常呈现一致回应。研究呼吁通过更包容的训练方法缓解文化与人口群体偏见。

原文摘要 · Abstract (English)

This study evaluates the ability of DeepSeek, an open-source large language model (LLM), to simulate public opinions in comparison to LLMs developed by major tech companies. By comparing DeepSeek-R1 and DeepSeek-V3 with Qwen2.5, GPT-4o, and Llama-3.3 and utilizing survey data from the American National Election Studies (ANES) and the Zuobiao dataset of China, we assess these models' capacity to predict public opinions on social issues in both China and the United States, highlighting their comparative capabilities between countries. Our findings indicate that DeepSeek-V3 performs best in simulating U.S. opinions on the abortion issue compared to other topics such as climate change, gun control, immigration, and services for same-sex couples, primarily because it more accurately simulates responses when provided with Democratic or liberal personas. For Chinese samples, DeepSeek-V3 performs best in simulating opinions on foreign aid and individualism but shows limitations in modeling views on capitalism, particularly failing to capture the stances of low-income and non-college-educated individuals. It does not exhibit significant differences from other models in simulating opinions on traditionalism and the free market. Further analysis reveals that all LLMs exhibit the tendency to overgeneralize a single perspective within demographic groups, often defaulting to consistent responses within groups. These findings highlight the need to mitigate cultural and demographic biases in LLM-driven public opinion modeling, calling for approaches such as more inclusive training methodologies.

大模型民意模拟跨文化偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。