测试大模型能否真实模拟人类性格在纠纷中的行为差异
Can LLMs Truly Embody Human Personality? Analyzing AI and Human Behavior Alignment in Dispute Resolution
- 构建对比框架,分析人类与大模型在纠纷对话中的人格表现
- 发现不同大模型对同一人格特质的冲突行为响应差异显著
- 提醒在社交应用中使用前需心理层面验证模型可靠性
大语言模型(LLMs)正被用于模拟法律调解、谈判和纠纷解决等社会场景中的行为。然而,这些模拟是否能再现人类个性与行为的关联尚不明确。人的性格会影响其在情绪化互动中的策略选择。本文提出一种评估框架,可直接比较人类-人类与大模型-大模型在纠纷对话中基于五大性格量表(BFI)的表现,提供可解释的战略行为与冲突结果指标。同时,我们开发了一种新型数据生成方法,创建了匹配场景与人格特质的大型模型纠纷对话数据集。通过对三个主流闭源大模型的应用,结果显示不同模型在人格驱动的冲突行为上存在显著差异,远超人类数据中的变异范围,挑战了‘人格提示代理可作为可靠行为替代’的假设。研究强调,人工智能在社会应用前必须进行心理学基础验证。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to simulate human behavior in social settings such as legal mediation, negotiation, and dispute resolution. However, it remains unclear whether these simulations reproduce the personality-behavior patterns observed in humans. Human personality, for instance, shapes how individuals navigate social interactions, including strategic choices and behaviors in emotionally charged interactions. This raises the question: Can LLMs, when prompted with personality traits, reproduce personality-driven differences in human conflict behavior? To explore this, we introduce an evaluation framework that enables direct comparison of human-human and LLM-LLM behaviors in dispute resolution dialogues with respect to Big Five Inventory (BFI) personality traits. This framework provides a set of interpretable metrics related to strategic behavior and conflict outcomes. We additionally contribute a novel dataset creation methodology for LLM dispute resolution dialogues with matched scenarios and personality traits with respect to human conversations. Finally, we demonstrate the use of our evaluation framework with three contemporary closed-source LLMs and show significant divergences in how personality manifests in conflict across different LLMs compared to human data, challenging the assumption that personality-prompted agents can serve as reliable behavioral proxies in socially impactful applications. Our work highlights the need for psychological grounding and validation in AI simulations before real-world use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。