arXiv:2412.13169cs.CL2024-12ACL被引 17

用大模型生成德国民意,发现不同模型表现差异大且有政治偏见。

Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case Study

  • 基于德国选民数据,用人口特征定制提示词生成虚拟民意。
  • Llama 在低多样性群体中表现最佳,对绿党等左翼支持者更准确。
  • 模型对极右翼政党AfD的模拟最差,变量选择影响显著。

近期研究中,大型语言模型(LLMs)被广泛用于探究公众意见。本研究探讨了LLMs的算法保真度,即其在再现人类参与者社会文化背景和细微观点方面的能力。基于德国纵向选举研究(GLES)的开放式调查数据,我们通过将人口统计特征融入角色提示词,让不同LLMs生成反映德国子群体的合成公众意见。结果显示,Llama在代表子群体方面优于其他模型,尤其在组内观点多样性较低时表现更佳。研究还发现,模型对绿色联盟和左翼党的支持者模拟效果更好,而与极右翼政党AfD的支持者匹配度最低。此外,提示词中包含或排除特定变量会显著影响模型预测结果。这些发现强调了调整模型以更有效建模多元公众意见、减少政治偏见并提升代表性鲁棒性的必要性。

原文摘要 · Abstract (English)

In recent research, large language models (LLMs) have been increasingly used to investigate public opinions. This study investigates the algorithmic fidelity of LLMs, i.e., the ability to replicate the socio-cultural context and nuanced opinions of human participants. Using open-ended survey data from the German Longitudinal Election Studies (GLES), we prompt different LLMs to generate synthetic public opinions reflective of German subpopulations by incorporating demographic features into the persona prompts. Our results show that Llama performs better than other LLMs at representing subpopulations, particularly when there is lower opinion diversity within those groups. Our findings further reveal that the LLM performs better for supporters of left-leaning parties like The Greens and The Left compared to other parties, and matches the least with the right-party AfD. Additionally, the inclusion or exclusion of specific variables in the prompts can significantly impact the models' predictions. These findings underscore the importance of aligning LLMs to more effectively model diverse public opinions while minimizing political biases and enhancing robustness in representativeness.

大模型民意生成偏见分析德国数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。