通过多轮交互评估大模型的人格化行为,发现多数行为需多次对话才出现。
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
- 设计多轮评测框架,覆盖14种人格化行为
- 1101名用户参与验证,模型行为与真实感知高度相关
- 揭示大模型人格化行为依赖持续互动,适合伦理与交互研究者
用户将大语言模型(LLMs)人格化的倾向日益引起开发者、研究者和政策制定者的关注。本文提出一种新方法,用于在真实多样场景中实证评估大模型的人格化行为。超越单轮静态基准,我们实现三项方法论突破:首先,构建包含14种人格化行为的多轮评估体系;其次,采用模拟用户交互的可扩展自动化流程;第三,开展大规模人机交互实验(N=1101),验证所测行为能预测真实用户的拟人感知。结果表明,所有SOTA LLM均表现出相似特征,如建立关系(如共情与认可)和第一人称代词使用,且多数行为仅在多轮对话后才显现。本研究为探究设计选择如何影响人格化行为提供了实证基础,并推动了相关伦理讨论。同时凸显了对复杂人机社交现象进行多轮评估的必要性。
原文摘要 · Abstract (English)
The tendency of users to anthropomorphise large language models (LLMs) is of growing interest to AI developers, researchers, and policy-makers. Here, we present a novel method for empirically evaluating anthropomorphic LLM behaviours in realistic and varied settings. Going beyond single-turn static benchmarks, we contribute three methodological advances in state-of-the-art (SOTA) LLM evaluation. First, we develop a multi-turn evaluation of 14 anthropomorphic behaviours. Second, we present a scalable, automated approach by employing simulations of user interactions. Third, we conduct an interactive, large-scale human subject study (N=1101) to validate that the model behaviours we measure predict real users' anthropomorphic perceptions. We find that all SOTA LLMs evaluated exhibit similar behaviours, characterised by relationship-building (e.g., empathy and validation) and first-person pronoun use, and that the majority of behaviours only first occur after multiple turns. Our work lays an empirical foundation for investigating how design choices influence anthropomorphic model behaviours and for progressing the ethical debate on the desirability of these behaviours. It also showcases the necessity of multi-turn evaluations for complex social phenomena in human-AI interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。