arXiv:2602.23964cs.IR2026-02被引 3

解决电商生成式检索中偏好对齐难题,提升搜索精度与训练效率

RAD-DPO: Robust Adaptive Denoising Direct Preference Optimization for Generative Retrieval in E-commerce

  • 通过令牌级梯度解耦保护语义前缀结构,避免梯度冲突
  • 引入动态奖励加权机制,降低隐式反馈噪声影响
  • 多标签全局对比损失扩大正样本覆盖,适合工业级部署

生成式检索(GR)正快速重塑电商搜索,以自回归解码结构化语义标识符(SIDs)替代传统多阶段流程。尽管架构高效,但将GR模型与真实用户偏好对齐仍是关键挑战。直接应用直接偏好优化(DPO)于结构化SIDs存在三方面问题:(i) 损害共享层级前缀,引发梯度冲突;(ii) 对隐式反馈产生的伪负样本敏感;(iii) 在多标签查询中加剧有效候选项间的概率“挤压效应”。为此,我们提出RAD-DPO,引入令牌级梯度解耦以保护前缀结构,采用基于相似度的动态奖励加权缓解标签噪声,并结合全局对比目标与全局SFT损失,显式扩展正样本覆盖范围。在京东核心搜索引擎上开展的大规模离线评估与在线A/B测试表明,RAD-DPO显著提升检索精度与训练效率,验证其在大规模工业部署中的鲁棒性。

原文摘要 · Abstract (English)

Generative Retrieval (GR) is rapidly transforming e-commerce search by replacing traditional multi-stage pipelines with the autoregressive decoding of structured Semantic IDs (SIDs). Despite this architectural efficiency, aligning GR models with nuanced, real-world user preferences remains a critical challenge. While Direct Preference Optimization (DPO) offers an efficient alignment solution, its direct application to structured SIDs suffers from three limitations: (i) it penalizes shared hierarchical prefixes, causing gradient conflicts; (ii) it is vulnerable to noisy pseudo-negatives from implicit feedback; and (iii) in multi-label queries with multiple relevant items, it exacerbates a probability "squeezing effect" among valid candidates. To address these issues, we propose RAD-DPO, which introduces token-level gradient detachment to protect prefix structures, similarity-based dynamic reward weighting to mitigate label noise, and a multi-label global contrastive objective integrated with global SFT loss to explicitly expand positive coverage. Extensive offline evaluations and large-scale online A/B testing on JD.com's core search engine demonstrate that RAD-DPO achieves significant improvements in both retrieval precision and training efficiency, proving its robustness for massive industrial deployments.

生成式检索偏好对齐电商搜索DPO改进

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。