arXiv:2605.25758cs.CL2026-05

构建实时用户画像评估基准,揭示大模型在动态兴趣变化中的滞后问题。

StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios

论文配图:StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios
图 1 · 摘自论文原文
  • 将用户画像建模为持续更新的状态维护任务,利用兴趣时间相关性设计无标注评估框架。
  • 基于7000+真实用户12万条内容的跨平台数据,测试14个主流大模型,发现普遍存在保守偏差。
  • 适合关注个性化推荐、动态用户建模的研究者与工业界开发者参考。

大语言模型重塑了用户画像技术,但现有评估多聚焦静态数据快照,忽视了个性化系统中用户生成内容(UGC)持续流入、画像快速演化的现实。为此,我们提出StreamProfileBench,一个大规模的细粒度流式用户画像评估基准。我们将流式用户画像形式化为连续状态维护任务,并构建了一个包含超过12万条来自五个多样化平台的7000+真实用户UGC数据集。通过利用用户兴趣的时间相关性,我们进一步提出一种无需标注的新型评估框架。在14个主流大语言模型上的大量实验表明,持续画像更新仍是一个开放挑战:模型普遍存在系统性保守偏差,过度保留过往兴趣而未能识别兴趣衰减。消融实验进一步验证了流式范式的实用价值与必要性。数据与代码已开源至https://github.com/WaterWang-001/StreamProfileBench。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the reality of personalized systems, where User-Generated Content (UGC) arrives continuously and fine-grained profiles evolve rapidly. To bridge this gap, we introduce StreamProfileBench, a large-scale benchmark for fine-grained streaming user profiling. We formalize streaming user profiling as a continuous state maintenance task and curate a highly authentic dataset comprising over 120,000 UGC posts from 7,000+ real users across five diverse platforms. By leveraging the temporal correlation of user interests, we further propose a novel, annotation-free evaluation framework. Extensive experiments across 14 leading LLMs reveal that continuous profile updating remains an open challenge. Models exhibit a systemic conservative bias, over-retaining past interests while failing to recognize interest decay. Ablation experiments further validate the practical utility and necessity of the streaming paradigm. Data and code are hosted in https://github.com/WaterWang-001/StreamProfileBench.

用户画像流式数据大模型评估兴趣演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。