arXiv:2605.07632cs.CLcs.AI2026-05被引 3

后训练让大模型更不像人,反而降低对人类行为的模拟精度。

Post-training makes large language models less human-like

论文配图:Post-training makes large language models less human-like
图 1 · 摘自论文原文
  • 通过新数据集Psych-201量化模型与人类行为的对齐度。
  • 后训练阶段普遍降低模型与人类行为的一致性,且新模型更严重。
  • 用个性化提示也难以提升个体预测效果,适合做行为建模研究者关注。

大型语言模型(LLMs)越来越多地被用作人类参与者的替代品,但尚不清楚哪些模型最能捕捉人类行为及其原因。为解决这一问题,我们引入了Psych-201,一个新型数据集,可实现大规模行为对齐度量。研究发现,后训练——将基础模型转变为有用助手的阶段——在不同模型家族、规模和目标下均一致降低了与人类行为的对齐度。此外,尽管基础模型持续改进,新模型世代中的这种偏差仍在扩大。最后,我们发现人格诱导(persona-induction)——一种通过条件化特定参与者信息来激发类人行为的常用技术——并未在个体层面提升预测性能。综上,当前用于将大模型转化为有用助手的过程,反而使其成为人类行为的更差模型。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as surrogates for human participants, but it remains unclear which models best capture human behavior and why. To address this, we introduce Psych-201, a novel dataset that enables us to measure behavioral alignment at scale. We find that post-training -- the stage that turns base models into useful assistants -- consistently reduces alignment with human behavior across model families, sizes, and objectives. Moreover, this misalignment widens in newer model generations even as base models continue to improve. Finally, we find that persona-induction -- a popular technique for eliciting human-like behavior by conditioning models on participant-specific information -- does not improve predictions at the level of individuals. Taken together, our results suggest that the very processes that are currently employed to turn LLMs into useful assistants also make them less accurate models of human behavior.

大模型行为对齐后训练心理学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。