arXiv:2509.04751cs.IRcs.LG2025-09ICML被引 2

用多模态大模型更准预测用户短视频兴趣变化。

Multimodal Foundation Model-Driven User Interest Modeling and Behavior Analysis on Short Video Platforms

  • 融合视频、文字、音乐构建统一语义空间建模兴趣
  • 冷启动用户兴趣预测准确率提升显著,点击率更高
  • 可解释性设计让推荐逻辑更透明,适合平台优化

随着短视频平台用户规模快速扩大,个性化推荐系统在提升用户体验和优化内容分发中作用日益关键。传统兴趣建模方法多依赖单一模态数据(如点击日志或文本标签),难以在复杂多模态内容环境中全面捕捉用户偏好。为此,本文提出一种基于多模态基础模型的用户兴趣建模与行为分析框架。通过交叉模态对齐策略,将视频帧、文本描述与背景音乐整合至统一语义空间,构建细粒度用户兴趣向量。同时引入行为驱动的特征嵌入机制,结合观看、点赞、评论序列建模兴趣动态演化,显著提升推荐的时效性与准确性。实验在公开与私有短视频数据集上进行,对比多种主流推荐算法与建模方法,结果表明:本方法在行为预测准确率、冷启动用户兴趣建模及推荐点击率方面均有显著提升。此外,通过注意力权重与特征可视化实现可解释性分析,揭示多模态输入下的决策依据并追踪兴趣演变路径,增强了推荐系统的透明性与可控性。

原文摘要 · Abstract (English)

With the rapid expansion of user bases on short video platforms, personalized recommendation systems are playing an increasingly critical role in enhancing user experience and optimizing content distribution. Traditional interest modeling methods often rely on unimodal data, such as click logs or text labels, which limits their ability to fully capture user preferences in a complex multimodal content environment. To address this challenge, this paper proposes a multimodal foundation model-based framework for user interest modeling and behavior analysis. By integrating video frames, textual descriptions, and background music into a unified semantic space using cross-modal alignment strategies, the framework constructs fine-grained user interest vectors. Additionally, we introduce a behavior-driven feature embedding mechanism that incorporates viewing, liking, and commenting sequences to model dynamic interest evolution, thereby improving both the timeliness and accuracy of recommendations. In the experimental phase, we conduct extensive evaluations using both public and proprietary short video datasets, comparing our approach against multiple mainstream recommendation algorithms and modeling techniques. Results demonstrate significant improvements in behavior prediction accuracy, interest modeling for cold-start users, and recommendation click-through rates. Moreover, we incorporate interpretability mechanisms using attention weights and feature visualization to reveal the model's decision basis under multimodal inputs and trace interest shifts, thereby enhancing the transparency and controllability of the recommendation system.

多模态兴趣建模推荐系统可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。