arXiv:2601.10198cs.CL2026-01ACL

用心理模式构建人类行为基准,让大模型更像真人。

HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns

  • 将12000篇论文提炼为244种心理模式,模拟多模式交互
  • 80亿参数模型在多模式动态上超越320亿参数模型
  • 适合研究角色扮演、心理建模与可信人机交互的学者

大型语言模型在推理与生成方面表现卓越,是高级人格模拟和角色扮演语言代理的基础。然而,真正匹配人类认知与行为模式仍是关键挑战。我们提出HumanLLM框架,将心理模式视为相互作用的因果力。从约1.2万篇学术论文中构建244种模式,并合成11359个场景,其中2-5种模式相互强化、冲突或调节,通过多轮对话呈现内心想法、行为与对话。双层级检查表评估个体模式保真度与多模式动态涌现性,实现强人类对齐(r=0.90),揭示整体指标混淆了模拟准确性与社会宜人性。HumanLLM-8B虽仅含4倍少参数,但在多模式动态上优于Qwen3-32B,表明真实拟人化需认知建模——不仅要模拟人类行为,更要还原其心理生成机制。数据集、代码与模型已开源。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning and generation, serving as the foundation for advanced persona simulation and Role-Playing Language Agents (RPLAs). However, achieving authentic alignment with human cognitive and behavioral patterns remains a critical challenge for these agents. We present HumanLLM, a framework treating psychological patterns as interacting causal forces. We construct 244 patterns from $\sim$12,000 academic papers and synthesize 11,359 scenarios where 2-5 patterns reinforce, conflict, or modulate each other, with multi-turn conversations expressing inner thoughts, actions, and dialogue. Our dual-level checklists evaluate both individual pattern fidelity and emergent multi-pattern dynamics, achieving strong human alignment ($r=0.90$) while revealing that holistic metrics conflate simulation accuracy with social desirability. HumanLLM-8B outperforms Qwen3-32B on multi-pattern dynamics despite 4$\times$ fewer parameters, demonstrating that authentic anthropomorphism requires cognitive modeling -- simulating not just what humans do, but the psychological processes generating those behaviors. Our dataset, code, and model are available at:https://github.com/YJGoodbye2024/HumanLLM

心理建模角色扮演拟人化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。