首个评测大模型模拟人类连续行为的基准,揭示其与数字分身仍有差距。
How Far are LLMs from Being Our Digital Twins? A Benchmark for Persona-Based Behavior Chain Simulation
- 构建1001个角色的15846条行为链,包含详细背景信息。
- 顶尖模型在动态场景中连续行为预测准确率仍不理想。
- 适合研究数字孪生、个性化代理和行为建模的学者参考。
近期,大语言模型因其作为人类数字孪生的潜力而受到学术界广泛关注,这类虚拟代理旨在复制个体并自主完成决策、问题解决和推理等任务。然而,当前对大模型的评估多聚焦于对话模拟,忽视了对人类行为模拟的重要性。为此,我们提出了BehaviorChain——首个用于评估大模型持续人类行为模拟能力的基准。BehaviorChain包含1,001个独特人物的15,846条高质量、基于角色的行为链,每条行为链均附带详细的个人历史与元数据。评估时,我们将人物元数据融入大模型,并使其在BehaviorChain提供的动态场景中迭代推断上下文相关的合理行为。全面评估结果表明,即使最先进的模型在准确模拟连续人类行为方面依然表现不佳。
原文摘要 · Abstract (English)
Recently, LLMs have garnered increasing attention across academic disciplines for their potential as human digital twins, virtual proxies designed to replicate individuals and autonomously perform tasks such as decision-making, problem-solving, and reasoning on their behalf. However, current evaluations of LLMs primarily emphasize dialogue simulation while overlooking human behavior simulation, which is crucial for digital twins. To address this gap, we introduce BehaviorChain, the first benchmark for evaluating LLMs' ability to simulate continuous human behavior. BehaviorChain comprises diverse, high-quality, persona-based behavior chains, totaling 15,846 distinct behaviors across 1,001 unique personas, each with detailed history and profile metadata. For evaluation, we integrate persona metadata into LLMs and employ them to iteratively infer contextually appropriate behaviors within dynamic scenarios provided by BehaviorChain. Comprehensive evaluation results demonstrated that even state-of-the-art models struggle with accurately simulating continuous human behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。