arXiv:2606.02798cs.AI2026-06

用真实用户行为数据构建决策预测基准,验证个性化模型效果。

BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces

论文配图:BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces
图 1 · 摘自论文原文
  • 基于公开预言机与链上记录重建2000个钱包的决策历史。
  • 包含14万条信念预测与148万条交易预测实例,支持检索评估。
  • 揭示个性化提升信念预测效果,不同接口暴露模型不同缺陷。

许多决策支持场景需要适应个体用户的系统,但相关评估数据仍有限。现有用户理解基准多依赖模拟用户或模型生成行为,而近期研究指出模型模拟可能系统性偏离人类行为。我们提出 extsc{BehaviorBench},一个基于真实世界行为痕迹的个性化决策建模评估基准。该基准从公开预言机与链上记录中重构钱包级决策历史,组织为两个互补任务层: extit{信念预测}(预测用户在市场中的最终立场与信心)和 extit{交易预测}(预测单个交易的方向与金额)。在2000个评估钱包中,基准包含141,445个信念实例和1,485,972个交易实例,支持检索式评估的独立支撑池。我们在四种历史接口下评估前沿及开源生成模型:无个性化、直接近期历史、生成用户画像、检索支持钱包证据。个性化对信念预测的提升更一致,模型排名随任务层与指标变化,不同接口揭示不同失败模式。 extsc{BehaviorBench} 为研究个性化方法能否利用真实行为证据而非仅依赖模拟用户提供了评估场景。

原文摘要 · Abstract (English)

Many decision-support settings require systems that adapt to individual users, but evaluation data for this problem remain limited. Existing benchmarks for user understanding often rely on simulated users or model-generated behavior, even though recent work cautions that model-based simulations can diverge systematically from human behavior. We introduce \textsc{BehaviorBench}, a benchmark for evaluating personalized decision modeling from real-world behavioral traces. \textsc{BehaviorBench} reconstructs wallet-level decision histories from observed public prediction-market and on-chain records, and organizes them into two complementary task layers: \emph{Belief prediction}, which predicts a user's final revealed stance and confidence in a market, and \emph{Trade prediction}, which predicts the direction and amount of individual transactions. Across 2,000 evaluation wallets, the benchmark contains 141,445 Belief instances and 1,485,972 Trade instances, with disjoint support pools for retrieval-based evaluation. We evaluate frontier and open-weight generative models under four history interfaces: no personalization, direct recent history, generated user profiles, and retrieved support-wallet evidence. Personalization improves Belief prediction more consistently than Trade prediction, model rankings change across task layers and metrics, and different history interfaces expose different failure modes. \textsc{BehaviorBench} provides an evaluation setting for studying whether personalized methods can use real-world behavioral evidence rather than simulated users alone.

决策建模真实行为个性化链上数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。