arXiv:2510.16925cs.IR2025-10中稿 · WWW'26被引 6

让电商搜索更懂用户上下文,提升推荐精准度。

Towards Context-aware Reasoning-enhanced Generative Searching in E-commerce

  • 将用户行为、时空等多源上下文统一为文本表示,增强语义对齐。
  • 通过自进化训练,逐步提升模型推理能力,显著改善排序效果。
  • 针对强化学习中的偏差问题提出改进算法,适合高阶推荐系统研发者。

基于搜索的推荐是电商平台的核心应用场景。用户的复杂搜索上下文——如时空因素、历史交互与当前查询信息——构成了其决策的重要部分,反映了隐式偏好,补充了显式查询词。如何建模这些丰富上下文信号及其与候选商品的复杂关联仍是关键挑战。尽管已有大量研究致力于构建更有效的搜索方法,现有方案在整合上下文信息方面仍存在局限,难以充分捕捉用户意图。为此,我们提出一种上下文感知的推理增强生成搜索框架,以更好理解复杂上下文。该框架首先将异构的用户与商品上下文统一为文本表示或基于语义的标识符并进行对齐。为克服缺乏显式推理路径的问题,引入一种自进化后训练范式,通过迭代结合监督微调与强化学习,逐步增强模型推理能力。此外,我们识别出现有强化学习算法在搜索场景下的潜在偏差,并提出一种去偏化的GRPO变体以提升排序性能。在真实电商平台搜索日志数据上的大量实验表明,本方法优于多个强基线,验证了其有效性。

原文摘要 · Abstract (English)

Search-based recommendation is one of the most critical application scenarios in e-commerce platforms. Users' complex search contexts--such as spatiotemporal factors, historical interactions, and current query's information--constitute an essential part of their decision-making, reflecting implicit preferences that complement explicit query terms. Modeling such rich contextual signals and their intricate associations with candidate items remains a key challenge. Although numerous efforts have been devoted to building more effective search methods, existing approaches still show limitations in integrating contextual information, which hinders their ability to fully capture user intent. To address these challenges, we propose a context-aware reasoning-enhanced generative search framework for better \textbf{understanding the complicated context}. Specifically, the framework first unifies heterogeneous user and item contexts into textual representations or text-based semantic identifiers and aligns them. To overcome the lack of explicit reasoning trajectories, we introduce a self-evolving post-training paradigm that iteratively combines supervised fine-tuning and reinforcement learning to progressively enhance the model's reasoning capability. In addition, we identify potential biases in existing RL algorithms when applied to search scenarios and present a debiased variant of GRPO to improve ranking performance. Extensive experiments on search log data collected from a real-world e-commerce platform demonstrate that our approach achieves superior performance compared with strong baselines, validating its effectiveness for search-based recommendation.

电商搜索上下文建模推理增强强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。