LLM生成的人格在低资源环境下与真实人类感知差异大,尤其在共情和可信度上。
Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings
- 用文化相关问题对比人类与8个LLM人格的回应表现。
- 人类在所有感知维度上显著优于LLM,共情与可信度差距最大。
- LLM普遍存在乐观倾向(平均正向情感5.99),偏离真实人类水平。
近期进展使大型语言模型(LLMs)能够生成人工智能人格,但其缺乏深层情境、文化和情感理解,存在显著局限。本研究在孟加拉国这一低资源环境中,通过文化特定问题定量比较了人类反应与八种LLM生成的社会人格(如男性、女性、穆斯林、政治支持者)的响应表现。结果显示,人类在回答问题及所有人格感知矩阵中均显著优于所有LLM,尤其在共情与可信度方面差距明显。此外,LLM生成内容表现出系统性偏倚,符合“普莉安娜原则”,正向情感评分更高(LLMs平均Φ_{avg} = 5.99,人类为5.60)。这些发现表明,LLM人格未能准确反映资源匮乏环境中真实人群的体验。在将它们用于社会科学前,必须基于真实人类数据验证其一致性与可靠性。
原文摘要 · Abstract (English)
Recent advances enable Large Language Models (LLMs) to generate AI personas, yet their lack of deep contextual, cultural, and emotional understanding poses a significant limitation. This study quantitatively compared human responses with those of eight LLM-generated social personas (e.g., Male, Female, Muslim, Political Supporter) within a low-resource environment like Bangladesh, using culturally specific questions. Results show human responses significantly outperform all LLMs in answering questions, and across all matrices of persona perception, with particularly large gaps in empathy and credibility. Furthermore, LLM-generated content exhibited a systematic bias along the lines of the ``Pollyanna Principle'', scoring measurably higher in positive sentiment ($Φ_{avg} = 5.99$ for LLMs vs. $5.60$ for Humans). These findings suggest that LLM personas do not accurately reflect the authentic experience of real people in resource-scarce environments. It is essential to validate LLM personas against real-world human data to ensure their alignment and reliability before deploying them in social science research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。