提出新指标LIS,高效评估多模态推荐中模型对用户偏好的拟合上限。
A Metric for MLLM Alignment in Large-scale Recommendation
- 通过测量偏好数据泄露上限,间接评估MLLM与推荐系统的对齐程度。
- 在小红书内容流和展示广告上实测,用户停留时长和广告价值显著提升。
- 适合大规模部署多模态推荐模型的工程团队参考落地。
多模态推荐已成为现代推荐系统的关键技术,利用先进多模态大语言模型(MLLMs)的内容表征。为确保这些表征与推荐系统良好适配,其对齐性至关重要。然而,评估MLLM在推荐中的对齐性面临三大挑战:(1) 静态基准因真实场景动态性而不可靠;(2) 在线上系统中评估虽准确但规模化成本过高;(3) 传统指标在表征表现不佳时无法提供可操作洞察。为此,我们提出新型指标——泄漏影响得分(Leakage Impact Score, LIS),不直接评估MLLM,而是高效测量偏好数据的潜在上限。我们还分享了在真实场景中使用LIS部署MLLM的实用建议。在小红书探索页的内容流和展示广告的线上A/B测试中,该方法显著提升了用户停留时长和广告商价值。
原文摘要 · Abstract (English)
Multimodal recommendation has emerged as a critical technique in modern recommender systems, leveraging content representations from advanced multimodal large language models (MLLMs). To ensure these representations are well-adapted, alignment with the recommender system is essential. However, evaluating the alignment of MLLMs for recommendation presents significant challenges due to three key issues: (1) static benchmarks are inaccurate because of the dynamism in real-world applications, (2) evaluations with online system, while accurate, are prohibitively expensive at scale, and (3) conventional metrics fail to provide actionable insights when learned representations underperform. To address these challenges, we propose the Leakage Impact Score (LIS), a novel metric for multimodal recommendation. Rather than directly assessing MLLMs, LIS efficiently measures the upper bound of preference data. We also share practical insights on deploying MLLMs with LIS in real-world scenarios. Online A/B tests on both Content Feed and Display Ads of Xiaohongshu's Explore Feed production demonstrate the effectiveness of our proposed method, showing significant improvements in user spent time and advertiser value.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。