用大模型生成视频语义描述,让推荐更懂内容背后的故事和意图。
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
- 用现成多模态大模型自动生成视频自然语言描述
- 在MicroLens-100K上超越传统视觉/音频特征,提升推荐效果
- 无需微调,适配各类推荐系统,适合想提升理解力的推荐团队
现有视频推荐系统主要依赖用户定义的元数据或由专用编码器提取的低层视觉与声学信号。这些低层特征仅描述屏幕上出现的内容,而无法捕捉意图、幽默感和世界知识等深层语义,这些正是使片段引发观众共鸣的关键。例如,一个30秒的视频究竟是歌手在屋顶演唱,还是在土耳其卡帕多西亚奇岩地貌拍摄的讽刺性恶搞?此类差异对个性化推荐至关重要,但传统编码流程无法识别。本文提出一种简单、推荐系统无关的零微调框架,通过提示预训练多模态大模型(MLLM)将每个视频生成丰富的自然语言描述(如“带有滑稽打斗和管弦乐突刺的超级英雄恶搞”),从而弥合原始内容与用户意图之间的鸿沟。我们使用MLLM输出结果,结合先进文本编码器,并输入标准协同过滤、基于内容及生成式推荐模型。在模拟TikTok风格视频交互的MicroLens-100K数据集上,该框架在五种代表性模型中均优于传统的视频、音频和元数据特征。研究结果表明,利用MLLM作为实时知识提取器,有望构建更具意图感知能力的视频推荐系统。
原文摘要 · Abstract (English)
Existing video recommender systems rely primarily on user-defined metadata or on low-level visual and acoustic signals extracted by specialised encoders. These low-level features describe what appears on the screen but miss deeper semantics such as intent, humour, and world knowledge that make clips resonate with viewers. For example, is a 30-second clip simply a singer on a rooftop, or an ironic parody filmed amid the fairy chimneys of Cappadocia, Turkey? Such distinctions are critical to personalised recommendations yet remain invisible to traditional encoding pipelines. In this paper, we introduce a simple, recommendation system-agnostic zero-finetuning framework that injects high-level semantics into the recommendation pipeline by prompting an off-the-shelf Multimodal Large Language Model (MLLM) to summarise each clip into a rich natural-language description (e.g. "a superhero parody with slapstick fights and orchestral stabs"), bridging the gap between raw content and user intent. We use MLLM output with a state-of-the-art text encoder and feed it into standard collaborative, content-based, and generative recommenders. On the MicroLens-100K dataset, which emulates user interactions with TikTok-style videos, our framework consistently surpasses conventional video, audio, and metadata features in five representative models. Our findings highlight the promise of leveraging MLLMs as on-the-fly knowledge extractors to build more intent-aware video recommenders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。