用分步框架提升电商搜索相关性,三年累计提升27%效果
LORE: A Large Generative Model for Search Relevance
- 将相关性拆解为知识推理、多模态匹配、规则遵循三能力分别训练
- 通过两阶段训练和偏好对齐,实现线上好率指标累计提升27%
- 适合电商搜索、垂直领域大模型落地的从业者参考
我们提出LORE,一个面向电商搜索相关性的大型生成模型系统框架。经过三年迭代部署,该系统在在线好率(GoodRate)指标上实现了累计+27%的提升。本文总结了从数据、特征、训练、评估到部署全生命周期的经验。现有研究虽采用思维链(CoT)提升相关性,但常遇性能瓶颈,根源在于将相关性视为单一任务,缺乏系统分解。我们的核心洞见是:相关性由知识与推理、多模态匹配、规则遵从三种独立能力构成。为此,LORE提供完整的大模型相关性生命周期蓝图,包括:(1) 结合渐进式思维链合成与人类偏好对齐的两阶段训练范式;(2) 针对三大能力设计的综合评估基准RAIR;(3) 基于查询频率分层的部署策略,高效将离线模型能力迁移至线上系统。LORE既是实用解决方案,也为其他垂直领域提供方法论参考。
原文摘要 · Abstract (English)
Achievement. We introduce LORE, a systematic framework for Large Generative Model-based relevance in e-commerce search. Deployed and iterated over three years, LORE achieves a cumulative +27\% improvement in online GoodRate metrics. This report shares the valuable experience gained throughout its development lifecycle, spanning data, features, training, evaluation, and deployment. Insight. While existing works apply Chain-of-Thought (CoT) to enhance relevance, they often hit a performance ceiling. We argue this stems from treating relevance as a monolithic task, lacking principled deconstruction. Our key insight is that relevance comprises distinct capabilities: knowledge and reasoning, multi-modal matching, and rule adherence. We contend that a qualitative-driven decomposition is essential for breaking through current performance bottlenecks. Contributions. LORE provides a complete blueprint for the LLM relevance lifecycle. Key contributions include: (1) A two-stage training paradigm combining progressive CoT synthesis via SFT with human preference alignment via RL. (2) A comprehensive benchmark, RAIR, designed to evaluate these core capabilities. (3) A query frequency-stratified deployment strategy that efficiently transfers offline LLM capabilities to the online system. LORE serves as both a practical solution and a methodological reference for other vertical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。