测试大模型人格提示后是否像人一样决策,发现效果不理想。
Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans?
- 用经典心理实验测试人格提示后模型的社交行为
- 所有模型均出现行为失准,且提示微调无效
- 适合关注大模型社会性与可控性的研究者阅读
语言模型的快速发展催生了诸多新应用,部分依赖于大模型(LLMs)新兴的社会能力。如今,人们在人生关键时刻向虚拟伙伴寻求建议甚至倾诉秘密,凸显塑造模型人格的重要性。当前主流方法通过海量数据训练并施加人格提示以引导模型行为。本文探究两个核心问题:(i) 人格提示后的模型是否表现出与其人格相符的行为?(ii) 其行为能否被精细控制?我们采用米尔格拉姆实验和最后通牒博弈作为社交行为测试场景,对来自四个厂商的开源与闭源模型进行人格提示实验。结果揭示出所有测试模型均存在行为失准的共性缺陷,且在提示扰动下依然持续存在。这些发现挑战了社区对人格提示技术普遍乐观的看法。
原文摘要 · Abstract (English)
The ongoing revolution in language modeling has led to various novel applications, some of which rely on the emerging social abilities of large language models (LLMs). Already, many turn to the new cyber friends for advice during the pivotal moments of their lives and trust them with the deepest secrets, implying that accurate shaping of the LLM's personality is paramount. To this end, state-of-the-art approaches exploit a vast variety of training data, and prompt the model to adopt a particular personality. We ask (i) if personality-prompted models behave (i.e., make decisions when presented with a social situation) in line with the ascribed personality (ii) if their behavior can be finely controlled. We use classic psychological experiments, the Milgram experiment and the Ultimatum Game, as social interaction testbeds and apply personality prompting to open- and closed-source LLMs from 4 different vendors. Our experiments reveal failure modes of the prompt-based modulation of the models' behavior that are shared across all models tested and persist under prompt perturbations. These findings challenge the optimistic sentiment toward personality prompting generally held in the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。