arXiv:2608.04809cs.IR2026-08KDD

DEGR通过双探索机制提升推荐系统跨请求上下文关联能力

DEGR: Dual Exploration-Driven Generative Re-Ranking for Adaptive Cross-Request Context Bridging

论文配图:DEGR: Dual Exploration-Driven Generative Re-Ranking for Adaptive Cross-Request Context Bridging
图 1 · 摘自论文原文
  • 采用监督与强化学习融合的双探索优化框架
  • 在京东电商系统中实现最高1.22% UCTR提升
  • 适合处理低质量内容供给场景的推荐优化

在工业级推荐系统中,重排序阶段需在业务目标与多样性之间权衡,并建模上下文信息。然而,受限于固定的上游供给,现有方法在低质量供给下难以进一步提升效果。为此,我们提出双探索驱动生成式重排序(DEGR)方法。DEGR采用监督-强化学习融合的探索与优化范式,由探索奖励模型引导,自适应平衡即时价值与探索价值。该混合优化框架包含三个核心组件:监督学习、探索多样性约束,以及自适应奖励加权的ORPO偏好优化。通过双重探索,生成器最终充当自适应的跨请求上下文桥梁。离线与在线实验表明,DEGR优于当前最优方法,在京东电商平台实现最高1.22% UCTR和0.20% PV提升。

原文摘要 · Abstract (English)

In industrial recommendation systems, the re-ranking stage balances business objectives and diversity for sequence-level optimization while modeling contextual information. However, constrained by fixed upstream supply, existing methods fail to deliver further effectiveness gains, especially under low-quality supply. To overcome this, re-ranking can actively balance immediate and exploratory value, for instance, by prioritizing exploratory exposure under low-quality supply to preserve browsing potential and facilitate serendipitous conversions. Therefore, we propose a Dual Exploration-Driven Generative Re-Ranking (DEGR) method. DEGR adopts a hybrid supervised-reinforcement exploration and optimization paradigm, guided by an exploratory reward model that adaptively balances immediate and exploratory value. The hybrid optimization paradigm integrates three key components: supervised learning, exploration diversity constraint, and adaptive reward-weighted ORPO for preference optimization. Through this dual exploration, the generator ultimately acts as an adaptive cross-request contextual bridge. Offline and online experiments indicate that DEGR outperforms SOTA methods, achieving improvements of up to 1.22% UCTR and 0.20% PV in the JD E-commerce recommendation system.

推荐系统生成重排序探索机制电商推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。