用时间感知检索提升短视频热度预测准确率
M3TR: Temporal Retrieval Enhanced Multi-Modal Micro-video Popularity Prediction
- 融合用户互动时序特征与视频流行趋势,建模反馈动态
- 在真实数据集上比现有方法低19.3%的预测误差
- 适合做短视频推荐、内容运营和平台算法优化
精准预测短视频热度是一项关键但极具挑战的任务,其用户互动呈现剧烈波动的‘过山车’式变化。现有方法常因对用户反馈动态理解肤浅,忽略点赞、评论、分享等行为间的相互激发与衰减特性;同时依赖静态内容相似性进行检索,忽视视频热度随时间演变的关键模式。为此,我们提出M³TR——一种时间感知检索增强的多模态微视频热度预测框架。核心创新在于引入Mamba-Hawkes过程(MHP)模块,显式建模用户反馈为自激事件序列,捕捉互动中的长程依赖关系;并构建时间感知检索引擎,基于多模态内容(视觉、音频、文本)与流行轨迹双重相似性,检索历史相关视频。通过融合检索知识增强目标视频特征,实现更全面的预测理解。在两个真实数据集上的实验表明,M³TR达到当前最优性能,在nMSE指标上相比之前方法最高提升19.3%,显著改善长期预测效果。
原文摘要 · Abstract (English)
Accurately predicting the popularity of micro-videos is a critical but challenging task, characterized by volatile, `rollercoaster-like' engagement dynamics. Existing methods often fail to capture these complex temporal patterns, leading to inaccurate long-term forecasts. This failure stems from two fundamental limitations: \ding{172} a superficial understanding of user feedback dynamics, which overlooks the mutually exciting and decaying nature of interactions such as likes, comments, and shares; and~\ding{173} retrieval mechanisms that rely solely on static content similarity, ignoring the crucial patterns of how a video's popularity evolves over time. To address these limitations, we propose \textbf{M$^3$TR}, a \textbf{T}emporal \textbf{R}etrieval enhanced \textbf{M}ulti-\textbf{M}odal framework that uniquely synergizes fine-grained temporal modeling with a novel temporal-aware retrieval process for \textbf{M}icro-video popularity prediction. At its core, M$^3$TR introduces a Mamba-Hawkes Process (MHP) module to explicitly model user feedback as a sequence of self-exciting events, capturing the intricate, long-range dependencies within user interactions (for \textbf{limitation} \ding{172}). This rich temporal representation then powers a temporal-aware retrieval engine that identifies historically relevant videos based on a combined similarity of both their multi-modal content (visual, audio, text) and their popularity trajectories (for \textbf{limitation} \ding{173}). By augmenting the target video's features with this retrieved knowledge, M$^3$TR achieves a comprehensive understanding of prediction. Extensive experiments on two real-world datasets demonstrate the superiority of our framework. M$^3$TR achieves state-of-the-art performance, outperforming previous methods by up to \textbf{19.3}\% in nMSE and showing significant gains in addressing long-term prediction challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。