用分层扩散模型捕捉视频推荐中的多粒度时序偏好,提升推荐准确率。
MealRec: Multi-granularity Sequential Modeling via Hierarchical Diffusion Models for Micro-Video Recommendation
- 通过分层扩散模型建模视频内与跨视频的时序关系。
- 在四个平台数据集上显著优于基线方法,提升推荐效果。
- 适合做短视频推荐系统优化的研究者和工程师参考。
微视频推荐旨在从用户互动视频的协同与上下文信息中捕捉偏好,进而预测合适视频。然而,多模态内容中的固有噪声和不可靠的隐式反馈会削弱行为与潜在兴趣之间的对应关系。传统方法虽采用行为增强建模与内容中心的多模态分析,但易引发非相关视频表示提取和模态冲突问题。为此,我们提出基于分层扩散模型的多粒度序列建模方法(MealRec),同时考虑视频内与跨视频的时序关联。首先设计时序引导的内容扩散(TCD),在视频内部时序与个性化协同信号指导下优化视频表示,突出关键内容并抑制冗余。为进一步实现语义一致的偏好建模,提出无噪声条件的偏好去噪(NPD),在无监督去噪条件下恢复被污染的用户偏好。在两个平台的四个微视频数据集上的大量实验表明,MealRec具有显著有效性、普适性与鲁棒性,且验证了TCD与NPD的有效机制。代码与数据集将在录用后公开。
原文摘要 · Abstract (English)
Micro-video recommendation aims to capture user preferences from the collaborative and context information of the interacted micro-videos, thereby predicting the appropriate videos. This target is often hindered by the inherent noise within multimodal content and unreliable implicit feedback, which weakens the correspondence between behaviors and underlying interests. While conventional works have predominantly approached such scenario through behavior-augmented modeling and content-centric multimodal analysis, these paradigms can inadvertently give rise to two non-trivial challenges: preference-irrelative video representation extraction and inherent modality conflicts. To address these issues, we propose a Multi-granularity sequential modeling method via hierarchical diffusion models for micro-video Recommendation (MealRec), which simultaneously considers temporal correlations during preference modeling from intra- and inter-video perspectives. Specifically, we first propose Temporal-guided Content Diffusion (TCD) to refine video representations under intra-video temporal guidance and personalized collaborative signals to emphasize salient content while suppressing redundancy. To achieve the semantically coherent preference modeling, we further design the Noise-unconditional Preference Denoising (NPD) to recovers informative user preferences from corrupted states under the blind denoising. Extensive experiments and analyses on four micro-video datasets from two platforms demonstrate the effectiveness, universality, and robustness of our MealRec, further uncovering the effective mechanism of our proposed TCD and NPD. The source code and corresponding dataset will be available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。