arXiv:2501.15283cs.CL2025-01被引 2

测试大模型代理的称谓使用是否像人,发现差距明显。

Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions

  • 用大模型模拟领导与非领导者的对话,观察称谓差异
  • 模型虽懂人类称谓规律,却无法在交互中真实体现
  • 提醒实践者慎用此类模拟做决策

随着大语言模型(LLMs)能力提升,研究者越来越多地将其用于社会模拟。本文探讨了基于LLM的代理间互动是否能模拟人类行为,重点关注领导者与非领导者在称谓使用上的差异,检验模型能否在交互中表现出类人称谓模式。评估结果表明,基于提示或专用代理的仿真与人类称谓使用存在显著差异,即使模型理解人类称谓规律,也无法在实际交互中展现。研究揭示了当前基于LLM的社会模拟存在局限性,警示从业者在决策过程中谨慎使用此类模拟。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) advance in their capabilities, researchers have increasingly employed them for social simulation. In this paper, we investigate whether interactions among LLM agents resemble those of humans. Specifically, we focus on the pronoun usage difference between leaders and non-leaders, examining whether the simulation would lead to human-like pronoun usage patterns during the LLMs' interactions. Our evaluation reveals the significant discrepancies between LLM-based simulations and human pronoun usage, with prompt-based or specialized agents failing to demonstrate human-like pronoun usage patterns. In addition, we reveal that even if LLMs understand the human pronoun usage patterns, they fail to demonstrate them in the actual interaction process. Our study highlights the limitations of social simulations based on LLM agents, urging caution in using such social simulation in practitioners' decision-making process.

社会模拟大模型称谓使用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。