arXiv:2602.03056cs.IR2026-02被引 2

用自然语言建模长期行为,精准预测用户兴趣组合。

ALPBench: A Benchmark for Attribution-level Long-term Personal Behavior Understanding

  • 将用户历史行为转为自然语言序列,实现可解释个性化
  • 聚焦属性组合预测,评估长时序行为理解能力
  • 支持新物品评估,适合研究长期兴趣建模的场景

大语言模型在个性化推荐中展现出巨大潜力,但准确捕捉用户偏好仍是关键挑战。本文提出ALPBench,一个面向归因级长期个人行为理解的基准测试。与传统以物品为中心的评测不同,ALPBench通过预测用户感兴趣的属性组合进行评估,即使对新引入的物品也能提供真实标签。该基准基于长期历史行为建模偏好,而非用户显式请求,更反映持久兴趣。用户历史以自然语言序列表示,支持可解释、基于推理的个性化。通过聚焦属性组合预测任务,该基准能精细评估个性化能力,该任务因需捕捉多属性间复杂交互及长序列推理而极具挑战性。

原文摘要 · Abstract (English)

Recent advances in large language models have highlighted their potential for personalized recommendation, where accurately capturing user preferences remains a key challenge. Leveraging their strong reasoning and generalization capabilities, LLMs offer new opportunities for modeling long-term user behavior. To systematically evaluate this, we introduce ALPBench, a Benchmark for Attribution-level Long-term Personal Behavior Understanding. Unlike item-focused benchmarks, ALPBench predicts user-interested attribute combinations, enabling ground-truth evaluation even for newly introduced items. It models preferences from long-term historical behaviors rather than users' explicitly expressed requests, better reflecting enduring interests. User histories are represented as natural language sequences, allowing interpretable, reasoning-based personalization. ALPBench enables fine-grained evaluation of personalization by focusing on the prediction of attribute combinations task that remains highly challenging for current LLMs due to the need to capture complex interactions among multiple attributes and reason over long-term user behavior sequences.

个性化推荐长时行为属性组合LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。