arXiv:2604.17259cs.IRcs.AI2026-04ACL

构建首个跨域长时序用户行为基准,推动真实场景下通用用户建模研究。

HORIZON: A Benchmark for In-the-wild User Behaviour Modeling

论文配图:HORIZON: A Benchmark for In-the-wild User Behaviour Modeling
图 1 · 摘自论文原文
  • 基于亚马逊评论数据重构跨域长时序用户行为数据集
  • 覆盖5400万用户与3500万商品,支持多任务评估与泛化测试
  • 提出时间、序列长度、未知用户等新挑战,适合长期建模研究者

真实世界中的用户行为具有多样性、跨域性与长时序特征。现有用户建模基准多局限于单领域短会话与下一个物品预测,难以支撑鲁棒且通用的用户模型发展。本文提出HORIZON,一个从大规模亚马逊评论数据重构的跨域用户行为基准,涵盖5400万用户与3500万商品,支持模型预训练与真实环境下的评估。该基准在数据、任务与评估三方面进行革新,要求模型在跨域、跨用户、跨时间条件下实现泛化,超越传统单一领域的缺失正例预测。我们设计了时间泛化、序列长度变化、未知用户建模等新任务及评估指标,更贴近真实部署场景。对比主流序列推荐架构与基于LLM的基线,结果揭示当前方法与真实需求间的显著差距,确立了HORIZON作为面向时空鲁棒、跨域通用用户模型研究的基础平台。

原文摘要 · Abstract (English)

User behavior in the real world is diverse, cross-domain, and spans long time horizons. Existing user modeling benchmarks however remain narrow, focusing mainly on short sessions and next-item prediction within a single domain. Such limitations hinder progress toward robust and generalizable user models. We present HORIZON, a new benchmark that reformulates user modeling along three axes i.e. dataset, task, and evaluation. Built from a large-scale, cross-domain reformulation of Amazon Reviews, HORIZON covers 54M users and 35M items, enabling both pretraining and realistic evaluation of models in heterogeneous environments. Unlike prior benchmarks, it challenges models to generalize across domains, users, and time, moving beyond standard missing-positive prediction in the same domain. We propose new tasks and evaluation setups that better reflect real-world deployment scenarios. These include temporal generalization, sequence-length variation, and modeling unseen users, with metrics designed to assess general user behavior understanding rather than isolated next-item prediction. We benchmark popular sequential recommendation architectures alongside LLM-based baselines that leverage long-term interaction histories. Our results highlight the gap between current methods and the demands of real-world user modeling, while establishing HORIZON as a foundation for research on temporally robust, cross-domain, and general-purpose user models.

用户建模跨域学习长时序基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。