通过频谱分解分离视频静态语义与动态变化,提升短视频推荐精准度。
Preference Flow Matching with Spectral Factorization for Micro-video Recommendation

- 在时频域用可学习掩码分离视频的静态语义与动态特征
- 用户敏感度加权后注入结构化上下文,引导偏好生成轨迹
- 在4个数据集上超越当前最优模型22.65%,推理成本最低
短视频推荐旨在从历史互动和多模态视频内容中推断用户偏好,从而预测其下一个感兴趣的视频。然而,现有方法将帧序列压缩为单一整体表示,混淆了稳定的视觉语义与演化的动态特征,二者共同影响用户偏好。同时,基于扩散和流匹配的推荐器仅依赖粗粒度行为上下文进行生成,忽略了其内部时间结构对偏好的塑造作用。为此,我们提出PrismRec,一种带有频谱分解的偏好流匹配框架。类似于棱镜将白光分解为光谱,PrismRec设计了频谱语义分解(SSF),通过时序频域中的先验引导可学习频率掩码,从帧级表示中提取互补的静态语义与动态因子。随后,提出上下文校准偏好匹配(CPM),根据每个用户的特定敏感度加权,并将校准后的上下文作为结构化条件注入,引导匹配轨迹向目标表示逼近,使视频内容成为偏好形成的核心驱动力而非辅助信息。在两个平台的四个数据集上的实验表明,PrismRec相比最先进基线最高提升22.65%,且推理开销和峰值内存均最低。
原文摘要 · Abstract (English)
Micro-video recommendation aims to infer user preferences from historical interactions and multimodal video content, thereby identifying the next video of interest. However, prevailing methods compress frame sequences into a single holistic representation, entangling the stable visual semantics and the evolving dynamics that jointly shape user preferences. Meanwhile, diffusion- and flow matching-based recommenders condition their generation process solely on coarse behavioral context, leaving its internal temporal structure outside preference formation. We therefore propose PrismRec, a Preference Flow Matching framework with Spectral Factorization for Micro-video Recommendation. Analogous to a prism that disperses white light into its constituent spectrum, PrismRec devises Spectral Semantic Factorization (SSF) to derive complementary static semantic and dynamic factors from frame-level representations via a prior-guided learnable frequency mask in the temporal frequency domain. Then, it proposes Context-Calibrated Preference Matching (CPM) to weigh them with each user's specific sensitivity and inject the calibrated context as a structured condition to steer the matching trajectory toward the target representation, making video content as an intrinsic driver of preference formation rather than auxiliary side information. Experiments on four datasets from two platforms show that PrismRec surpasses the SOTA baseline by up to 22.65%, with the lowest inference cost and peak memory among the compared methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。