arXiv:2410.21359cs.CLcs.AI2024-10被引 11

测试大模型在博弈中是否像人一样利他,发现身份设定不等于行为相似。

Can Machines Think Like Humans? A Behavioral Evaluation of LLM Agents in Dictator Games

  • 用不同人格设定测试大模型在独裁者博弈中的利他行为。
  • 模型行为与人类差异大,且不同模型表现无固定规律。
  • 提示词和模型架构对利他行为影响显著,适合社会科学研究者参考。

随着基于大语言模型(LLM)的智能体日益融入人类社会,我们对其亲社会行为的理解仍显不足。本文(1)探究不同人格设定如何诱导LLM智能体的亲社会行为,并与人类行为进行基准对比;(2)引入社会科学方法评估LLM智能体的决策机制。研究考察了不同人格设定与实验框架对同一模型家族内、跨模型家族以及与人类在独裁者博弈中利他行为的影响。结果表明,仅赋予大模型类人身份并不能使其表现出类人行为。这些发现提示,大模型推理过程在独裁者博弈中并未一致展现人类决策的文本标记,其与人类行为的对齐程度在不同模型架构与提示设计间存在显著差异,且无明确规律。随着机器智能深度融入社会,“亲社会人工智能”成为慈善研究中一个前景广阔且亟需探索的方向。

原文摘要 · Abstract (English)

As Large Language Model (LLM)-based agents increasingly engage with human society, how well do we understand their prosocial behaviors? We (1) investigate how LLM agents' prosocial behaviors can be induced by different personas and benchmarked against human behaviors; and (2) introduce a social science approach to evaluate LLM agents' decision-making. We explored how different personas and experimental framings affect these AI agents' altruistic behavior in dictator games and compared their behaviors within the same LLM family, across various families, and with human behaviors. The findings reveal that merely assigning a human-like identity to LLMs does not produce human-like behaviors. These findings suggest that LLM agents' reasoning does not consistently exhibit textual markers of human decision-making in dictator games and that their alignment with human behavior varies substantially across model architectures and prompt formulations; even worse, such dependence does not follow a clear pattern. As society increasingly integrates machine intelligence, "Prosocial AI" emerges as a promising and urgent research direction in philanthropic studies.

大模型行为利他行为博弈实验社会智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。