arXiv:2604.08362cs.CLcs.AI2026-04被引 7

首个基于真实数据的长期跨场景行为模拟基准,揭示大模型在行为多样性上的系统性偏差。

Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces

论文配图:Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces
图 1 · 摘自论文原文
  • 构建首个融合真实世界长期、跨场景、异构行为的统一模拟基准
  • 发现大模型在长周期决策中表现停滞,且趋向于平均化和理想化个体
  • 适合关注行为模拟、人机交互与大模型偏见的研究者

大语言模型(LLMs)为通用用户模拟带来了希望,但现有基准仍局限于孤立场景、狭窄动作空间或合成数据,无法捕捉真实人类行为的整体性。为此,我们提出了OmniBehavior,首个完全基于真实世界数据构建的行为模拟基准,将长期、跨场景、异构行为模式整合到统一框架中。基于该基准,我们首次提供实证证据表明:以往孤立场景的数据存在视野局限,而真实决策依赖长期、跨场景的因果链条。对主流LLM的广泛评估显示,当前模型在模拟复杂行为时表现不佳,即使上下文窗口扩展也出现性能瓶颈。关键发现是,模拟行为与真实行为的系统性对比揭示了根本性结构偏差:LLMs倾向于收敛至‘平均积极人格’,表现出过度活跃、人格同质化和乌托邦倾向,导致个体差异与长尾行为丢失,指明未来高保真模拟研究的关键方向。

原文摘要 · Abstract (English)

The emergence of Large Language Models (LLMs) has illuminated the potential for a general-purpose user simulator. However, existing benchmarks remain constrained to isolated scenarios, narrow action spaces, or synthetic data, failing to capture the holistic nature of authentic human behavior. To bridge this gap, we introduce OmniBehavior, the first user simulation benchmark constructed entirely from real-world data, integrating long-horizon, cross-scenario, and heterogeneous behavioral patterns into a unified framework. Based on this benchmark, we first provide empirical evidence that previous datasets with isolated scenarios suffer from tunnel vision, whereas real-world decision-making relies on long-term, cross-scenario causal chains. Extensive evaluations of state-of-the-art LLMs reveal that current models struggle to accurately simulate these complex behaviors, with performance plateauing even as context windows expand. Crucially, a systematic comparison between simulated and authentic behaviors uncovers a fundamental structural bias: LLMs tend to converge toward a positive average person, exhibiting hyper-activity, persona homogenization, and a utopian bias. This results in the loss of individual differences and long-tail behaviors, highlighting critical directions for future high-fidelity simulation research.

行为模拟大模型偏见真实数据长序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。