用多模态大模型提升推荐系统的意外惊喜感。
Serendipitous Recommendation with Multimodal LLM
- 用微调的多模态大模型为传统推荐系统提供高层指导,引导发现新颖内容。
- 在亿级视频平台实测,显著提升推荐惊喜度和用户满意度。
- 通过思维链策略挖掘用户未探索的兴趣簇,适合追求创新推荐的场景。
传统推荐系统擅长识别相关内容,但难以提供令人惊喜的新颖项目。多模态大语言模型(MLLM)具备世界知识和多模态理解能力,有利于实现惊喜推荐,但将其集成到亿级物品规模的平台仍面临挑战。本文提出一种分层框架:微调后的MLLM为传统推荐模型提供高层指引,引导其生成更具惊喜性的推荐。该方法利用MLLM对多模态内容和用户兴趣的理解优势,同时保留传统模型在物品级推荐中的高效性,降低直接将MLLM应用于庞大动作空间的复杂度。我们还展示了一种思维链策略,使MLLM先理解视频内容,再识别相关但未探索的兴趣聚类,从而发现新兴趣。在服务于数十亿用户的商业短视频平台进行的在线实验表明,该MLLM驱动的方法显著提升了推荐的惊喜度与用户满意度。
原文摘要 · Abstract (English)
Conventional recommendation systems succeed in identifying relevant content but often fail to provide users with surprising or novel items. Multimodal Large Language Models (MLLMs) possess the world knowledge and multimodal understanding needed for serendipity, but their integration into billion-item-scale platforms presents significant challenges. In this paper, we propose a novel hierarchical framework where fine-tuned MLLMs provide high-level guidance to conventional recommendation models, steering them towards more serendipitous suggestions. This approach leverages MLLM strengths in understanding multimodal content and user interests while retaining the efficiency of traditional models for item-level recommendation. This mitigates the complexity of applying MLLMs directly to vast action spaces. We also demonstrate a chain-of-thought strategy enabling MLLMs to discover novel user interests by first understanding video content and then identifying relevant yet unexplored interest clusters. Through live experiments within a commercial short-form video platform serving billions of users, we show that our MLLM-powered approach significantly improves both recommendation serendipity and user satisfaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。