arXiv:2507.20185cs.CL2025-07ACL被引 7

构建电商用户会话意图迁移基准,提升模型理解用户行为变化能力

SessionIntentBench: A Multi-task Inter-session Intention-shift Modeling Benchmark for E-commerce Customer Behavior Understanding

  • 提出意图树概念与数据构建流程,挖掘多会话间意图演变规律
  • 建立包含超195万条意图记录的基准,覆盖超100万会话轨迹
  • 验证大模型在复杂会话中捕捉意图迁移的能力不足,注入意图可提升性能

会话历史是记录用户在浏览过程中与多个商品交互行为的常见方式。例如,用户点击商品页面后离开,可能因某些特征不满足需求,这成为即时用户偏好的重要信号。然而,现有研究未能有效捕捉和建模用户意图,主要依赖标题、描述等表层信息,缺乏对深层意图的显式建模。同时,电商购买会话中专门用于意图建模的数据与基准也严重缺失。为此,本文引入意图树概念并设计数据构建流程,共同构建了多任务跨会话意图迁移基准 SessionIntentBench,评估大语言(视觉)模型在理解跨会话意图变化上的能力,包含四个子任务。该基准涵盖1,952,177条意图记录、1,132,145条会话意图轨迹,以及利用10,905个会话挖掘出的13,003,664个可用任务。通过人工标注获取部分数据的真实标签,形成评估金标准集。在标注数据上的大量实验表明,当前大语言(视觉)模型难以在复杂的会话场景中有效捕捉和利用意图变化。进一步分析显示,引入意图信息可显著提升大模型表现。

原文摘要 · Abstract (English)

Session history is a common way of recording user interacting behaviors throughout a browsing activity with multiple products. For example, if an user clicks a product webpage and then leaves, it might because there are certain features that don't satisfy the user, which serve as an important indicator of on-the-spot user preferences. However, all prior works fail to capture and model customer intention effectively because insufficient information exploitation and only apparent information like descriptions and titles are used. There is also a lack of data and corresponding benchmark for explicitly modeling intention in E-commerce product purchase sessions. To address these issues, we introduce the concept of an intention tree and propose a dataset curation pipeline. Together, we construct a sibling multimodal benchmark, SessionIntentBench, that evaluates L(V)LMs' capability on understanding inter-session intention shift with four subtasks. With 1,952,177 intention entries, 1,132,145 session intention trajectories, and 13,003,664 available tasks mined using 10,905 sessions, we provide a scalable way to exploit the existing session data for customer intention understanding. We conduct human annotations to collect ground-truth label for a subset of collected data to form an evaluation gold set. Extensive experiments on the annotated data further confirm that current L(V)LMs fail to capture and utilize the intention across the complex session setting. Further analysis show injecting intention enhances LLMs' performances.

意图建模电商推荐多会话分析大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。