arXiv:2411.13045cs.IRcs.AI2024-11中稿 · WWW 2025 oral被引 6

让大模型推理过程可解释,提升电商搜索相关性效果

Explainable LLM-driven Multi-dimensional Distillation for E-Commerce Relevance Learning

  • 用思维链分解大模型推理,提升可解释性
  • 同时蒸馏概率分布和推理过程,增强线上模型能力
  • 适合关注可解释性与搜索体验优化的从业者

有效的查询-商品相关性建模对提升电商平台搜索系统的用户体验至关重要。近年来,得益于其丰富的内在知识,大语言模型(LLM)在性能和长尾泛化能力上优于传统神经网络方法。然而,现有基于LLM的方法存在两大缺陷:一是参数量大、计算开销高,难以在线部署;二是相关性建模过程为黑盒,难以提取并应用其内部知识。为此,我们提出一种可解释的多维知识蒸馏框架(Explainable LLM-driven Multi-dimensional Distillation),包含两个核心组件:(1) 可解释的相关性建模大模型(ELLM-rele),将相关性学习分解为中间推理步骤,以思维链(CoT)方式建模,提升可解释性与性能;(2) 多维知识蒸馏(MKD)架构,从相关性得分分布与CoT推理两方面,将ELLM-rele的知识迁移至当前可部署的交互式与表示式学生模型。通过蒸馏概率与推理知识,MKD显著提升了学生模型的语义交互能力和长尾泛化能力。在淘宝搜索广告场景的离线评估与线上实验均表明,该框架显著提升了电商相关性建模性能与用户体验。

原文摘要 · Abstract (English)

Effective query-item relevance modeling is pivotal for enhancing user experience and safeguarding user satisfaction in e-commerce search systems. Recently, benefiting from the vast inherent knowledge, Large Language Model (LLM) approach demonstrates strong performance and long-tail generalization ability compared with previous neural-based specialized relevance learning methods. Though promising, current LLM-based methods encounter the following inadequacies in practice: First, the massive parameters and computational demands make it difficult to be deployed online. Second, distilling LLM models to online models is a feasible direction, but the LLM relevance modeling is a black box, and its rich intrinsic knowledge is difficult to extract and apply online. To improve the interpretability of LLM and boost the performance of online relevance models via LLM, we propose an Explainable LLM-driven Multi-dimensional Distillation framework for e-commerce relevance learning, which comprises two core components: (1) An Explainable LLM for relevance modeling (ELLM-rele), which decomposes the relevance learning into intermediate steps and models relevance learning as a Chain-of-Thought (CoT) reasoning, thereby enhancing both interpretability and performance of LLM. (2) A Multi-dimensional Knowledge Distillation (MKD) architecture that transfers the knowledge of ELLM-rele to current deployable interaction-based and representation-based student models from both the relevance score distribution and CoT reasoning aspects. Through distilling the probabilistic and CoT reasoning knowledge, MKD improves both the semantic interaction and long-tail generalization abilities of student models. Extensive offline evaluations and online experiments on Taobao search ad scene demonstrate that our proposed framework significantly enhances e-commerce relevance learning performance and user experience.

大模型知识蒸馏可解释性电商搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。