针对多场景推荐匹配阶段的冷门场景效果差问题,提出新模型提升稀疏场景召回质量。
Distillation-based Scenario-Adaptive Mixture-of-Experts for the Matching Stage of Multi-scenario Recommendation
- 用自适应投影生成轻量上下文参数,防止单一场景主导专家模型。
- 通过教师-学生知识蒸馏,让双塔模型学会捕捉复杂匹配模式。
- 在真实数据集上显著提升冷门场景的检索准确率,适合多场景推荐系统优化。
多场景推荐对优化跨场景用户体验至关重要。尽管多门混合专家(MMOE)在排序阶段表现优异,但其应用于匹配阶段受限于独立双塔架构的盲优化问题,以及头部场景参数主导带来的分布偏差。为此,我们提出基于知识蒸馏的场景自适应混合专家模型(DSMOE)。特别地,设计场景自适应投影(SAP)模块生成轻量化、上下文相关的参数,有效防止长尾场景中的专家崩溃。同时,引入跨架构知识蒸馏框架,由一个关注交互的教师模型指导双塔学生模型学习复杂匹配模式。在真实数据集上的大量实验表明,DSMOE在显著提升低频、数据稀疏场景的召回质量方面具有明显优势。
原文摘要 · Abstract (English)
Multi-scenario recommendation is pivotal for optimizing user experience across diverse contexts. While Multi-gate Mixture-of-Experts (MMOE) thrives in ranking, its transfer to the matching stage is hindered by the blind optimization inherent to independent two-tower architectures and the parameter dominance of head scenarios. To address these structural and distributional bottlenecks, we propose Distillation-based Scenario-Adaptive Mixture-of-Experts (DSMOE). Specially, we devise a Scenario-Adaptive Projection (SAP) module to generate lightweight, context-specific parameters, effectively preventing expert collapse in long-tail scenarios. Concurrently, we introduce a cross-architecture knowledge distillation framework, where an interaction-aware teacher guides the two-tower student to capture complex matching patterns. Extensive experiments on real-world datasets demonstrate DSMOE's superiority, particularly in significantly improving retrieval quality for under-represented, data-sparse scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。