arXiv:2507.19346cs.LG2025-07被引 1

用多模态模型提升电商短视频推荐,解决冷启动与位置偏差问题

Short-Form Video Recommendations with Multimodal Embeddings: Addressing Cold-Start and Bias Challenges

  • 基于微调的视觉-语言模型构建视频检索系统
  • 在线实验显示优于传统监督学习方法
  • 适合新视频内容上线时的冷启动场景

近年来,社交媒体用户在短视频平台花费大量时间。为吸引用户并延长停留时长,电商等其他领域平台也引入短视频内容。这类体验的成功不仅依赖内容本身,更得益于独特的用户界面设计:平台主动逐条推荐内容,而非提供可点击列表。这种沉浸式推荐流带来了新挑战,尤其在推出新视频功能时。除交互数据有限外,界面设计和观看时长优化还加剧了位置偏差,模型倾向于推荐更短的视频。这些因素与推荐系统的反馈循环共同导致效果难以提升。本文指出新短视频体验面临的核心挑战,并分享实践经验:即使具备充足的视频交互数据,使用微调后的多模态视觉-语言模型构建视频检索系统,仍比传统监督学习方法更具优势。在电商平台的在线实验中,该方案表现更优。

原文摘要 · Abstract (English)

In recent years, social media users have spent significant amounts of time on short-form video platforms. As a result, established platforms in other domains, such as e-commerce, have begun introducing short-form video content to engage users and increase their time spent on the platform. The success of these experiences is due not only to the content itself but also to a unique UI innovation: instead of offering users a list of choices to click, platforms actively recommend content for users to watch one at a time. This creates new challenges for recommender systems, especially when launching a new video experience. Beyond the limited interaction data, immersive feed experiences introduce stronger position bias due to the UI and duration bias when optimizing for watch-time, as models tend to favor shorter videos. These issues, together with the feedback loop inherent in recommender systems, make it difficult to build effective solutions. In this paper, we highlight the challenges faced when introducing a new short-form video experience and present our experience showing that, even with sufficient video interaction data, it can be more beneficial to leverage a video retrieval system using a fine-tuned multimodal vision-language model to overcome these challenges. This approach demonstrated greater effectiveness compared to conventional supervised learning methods in online experiments conducted on our e-commerce platform.

短视频推荐多模态嵌入冷启动位置偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。