arXiv:2410.03538cs.IRcs.AI2024-10

用多模态统一空间建模用户兴趣,提升短视频推荐精准度。

Dreaming User Multimodal Representation Guided by The Platonic Representation Hypothesis for Micro-Video Recommendation

  • 基于柏拉图表征假说,将用户行为映射到统一多模态空间。
  • 在线测试显示活跃天数与播放量显著提升,效果稳定。
  • 适合大规模短视频平台,尤其对冷启动用户有效。

在线微视频平台的兴起凸显了先进推荐系统在缓解信息过载、提供个性化内容方面的重要性。尽管已有进展,但准确且快速捕捉动态用户兴趣仍是重大挑战。受柏拉图表征假说启发——不同数据模态会收敛至对现实的共享统计模型——我们提出 DreamUMM(Dreaming User Multi-Modal Representation),一种利用用户历史行为在多模态空间中实时构建用户表征的新方法。DreamUMM 采用闭式解,将用户视频偏好与多模态相似性关联,假设用户兴趣可在统一多模态空间中有效表示。此外,针对缺乏近期行为数据的场景,我们提出 Candidate-DreamUMM,仅通过候选视频推断用户兴趣。大规模线上 A/B 测试表明,用户参与度指标(如活跃天数、播放次数)显著提升。DreamUMM 在两个日活超亿级的微视频平台成功部署,验证了其在个性化内容推送中的实际效能与可扩展性。本工作通过实证支持了用户兴趣表征可能存在于多模态空间中的假设,推动了表征收敛的探索。

原文摘要 · Abstract (English)

The proliferation of online micro-video platforms has underscored the necessity for advanced recommender systems to mitigate information overload and deliver tailored content. Despite advancements, accurately and promptly capturing dynamic user interests remains a formidable challenge. Inspired by the Platonic Representation Hypothesis, which posits that different data modalities converge towards a shared statistical model of reality, we introduce DreamUMM (Dreaming User Multi-Modal Representation), a novel approach leveraging user historical behaviors to create real-time user representation in a multimoda space. DreamUMM employs a closed-form solution correlating user video preferences with multimodal similarity, hypothesizing that user interests can be effectively represented in a unified multimodal space. Additionally, we propose Candidate-DreamUMM for scenarios lacking recent user behavior data, inferring interests from candidate videos alone. Extensive online A/B tests demonstrate significant improvements in user engagement metrics, including active days and play count. The successful deployment of DreamUMM in two micro-video platforms with hundreds of millions of daily active users, illustrates its practical efficacy and scalability in personalized micro-video content delivery. Our work contributes to the ongoing exploration of representational convergence by providing empirical evidence supporting the potential for user interest representations to reside in a multimodal space.

短视频推荐多模态表征用户建模冷启动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。