用分步审查机制提升电商搜索相关性,让大模型推理更准更可信。
SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance
- 分步奖励优化策略,结合生成模型与人工验证提供细粒度反馈
- 在真实电商数据上相关性准确率显著优于SFT、DPO等基线方法
- 适合需要高可解释性和鲁棒性的电商搜索系统开发者
查询-商品相关性预测对AI驱动的电商至关重要,但当前基于大模型的方法面临困境:SFT和DPO因粗粒度监督难以泛化长尾场景,传统RLVR则因反馈稀疏而无法修正中间推理错误。本文提出分步混合审查(SHE)强化学习框架,通过分步奖励策略优化(SRPO)确保逻辑一致性。SRPO采用生成式奖励模型与人工标注验证器相结合的混合奖励机制,提供细粒度的步骤级信号。为增强稳定性,SHE引入多样化数据过滤以维持策略熵,并设计多阶段课程学习协议实现渐进式技能提升。在真实电商搜索基准上的大量实验表明,SHE在大规模电商环境下显著提升了推理质量与相关性预测准确率,优于SFT、DPO、GRPO等基线方法,同时增强了可解释性与鲁棒性。
原文摘要 · Abstract (English)
Query-product relevance prediction is vital for AI-driven e-commerce, yet current LLM-based approaches face a dilemma: SFT and DPO struggle with long-tail generalization due to coarse supervision, while traditional RLVR suffers from sparse feedback that fails to correct intermediate reasoning errors. We propose Stepwise Hybrid Examination (SHE), an RL framework that ensures logical consistency through Stepwise Reward Policy Optimization (SRPO). SRPO utilizes a hybrid reward mechanism-combining generative reward models with human-annotated verifiers-to provide fine-grained, step-level signals. To further enhance stability, SHE incorporates diversified data filtering to maintain policy entropy and a multi-stage curriculum learning protocol for progressive skill acquisition. Extensive experiments on real-world search benchmarks show that SHE improves both reasoning quality and relevance-prediction accuracy in large-scale e-commerce settings, outperforming SFT, DPO, GRPO, and other baselines, while also enhancing interpretability and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。