分析大模型的人类行为表现,揭示其可控性与适用场景。
Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts

- 用大模型评分+人工评估,多维度测试4种模型的行为表现。
- 模型表现出普遍人类行为,但类型和程度随用户目标不同而异。
- 系统提示可调控行为,但需谨慎避免意外后果。
大型语言模型(LLMs)展现出丰富的人类行为特征,包括表达思想情感、与用户建立关系、拒绝请求并设定边界等。尽管这些行为普遍存在,研究者和从业者仍缺乏有效方法与实证洞察来判断何时以及何种行为应被启用。为此,我们通过“大模型作为裁判”和人工评估,开展多维度分析,考察这些行为的普遍性、潜在影响及可控性。基于来自四个主流模型(gpt-4o、gpt-4.1-mini、claude-sonnet-4.6、gemini-2.5-flash)的21,000轮多轮对话数据,发现人类行为在模型中广泛存在,且在不同模型及用户因素(对话目标与用户画像)下呈现差异。人工评估显示,自我指涉与关系构建行为在模型身上被认为比在人类身上更不恰当,而维持边界的行为则被认为在模型身上更合适。最后,我们证明系统提示可有效控制这些行为,但需经过仔细评估以避免意外影响。本文讨论了研究启示,并为负责任的大模型设计与评估提供建议。
原文摘要 · Abstract (English)
Large language models (LLMs) exhibit a wide range of human-like behaviors, from expressing thoughts and emotions, to engaging in relationship-building with users, to refusing requests and maintaining boundaries. Despite their prevalence, researchers and practitioners lack methods and empirical insights to make informed decisions about when and what types of human-like behaviors LLMs should exhibit. To fill this gap, we present a multi-dimensional analysis of the prevalence, potential effects, and controllability of these behaviors using LLM-as-a-judge and human evaluation. Across 21,000 multi-turn conversations from four widely used models (gpt-4o, gpt-4.1-mini, claude-sonnet-4.6, gemini-2.5-flash), we find that human-like behaviors are pervasive but vary across models and user factors (conversation goals and user profiles). In terms of perceived appropriateness, human evaluators judged self-referential and relationship-building behaviors as less appropriate from LLMs than from humans, but boundary-maintaining behaviors more appropriate from LLMs than from humans. Finally, we show that system prompting can control these behaviors, though it requires careful evaluation to avoid unintended effects. We discuss the implications of our findings and provide recommendations for responsible LLM design and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。