通过时空扩展提升短视频流行度预测精度。
Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction

- 构建时空联合扩展框架,融合稀疏采样与密集感知理解长视频内容。
- 在三个基准上超越11个基线,显著提升预测准确率与排序一致性。
- 适合关注短视频推荐与流量分配的工程师与研究者。
短视频流行度预测(MVPP)旨在预估在线媒体上视频未来的受欢迎程度,对内容推荐和流量分配至关重要。现实中,模型需同时理解视频自身的时间动态(时序)及其与历史视频的空间关联(空间)。现有方法在两方面均有局限:时序上依赖稀疏短程采样,限制内容感知;空间上依赖扁平检索记忆,容量有限且效率低,难以实现可扩展的知识利用。为此,本文提出统一框架,实现联合时空扩展,既能精准感知极长视频序列,又能支持可无限扩展的存储高效记忆库。技术上,采用由帧评分模块驱动的时序扩展,通过稀疏采样与密集感知两条互补路径提取关键帧线索,自适应融合以实现鲁棒的长序列理解。空间扩展则构建拓扑感知记忆库,基于拓扑关系分层聚类历史相关内容;新视频加入时仅更新对应簇的编码特征,实现无界历史关联而无需无界存储增长。在三个主流MVPP基准上的大量实验表明,本方法在主流指标上持续优于11个强基线,预测准确率与排序一致性均获稳健提升。
原文摘要 · Abstract (English)
Micro-video popularity prediction (MVPP) aims to forecast the future popularity of videos on online media, which is essential for applications such as content recommendation and traffic allocation. In real-world scenarios, it is critical for MVPP approaches to understand both the temporal dynamics of a given video (temporal) and its historical relevance to other videos (spatial). However, existing approaches sufer from limitations in both dimensions: temporally, they rely on sparse short-range sampling that restricts content perception; spatially, they depend on flat retrieval memory with limited capacity and low efficiency, hindering scalable knowledge utilization. To overcome these limitations, we propose a unified framework that achieves joint spatio-temporal enlargement, enabling precise perception of extremely long video sequences while supporting a scalable memory bank that can infinitely expand to incorporate all relevant historical videos. Technically, we employ a Temporal Enlargement driven by a frame scoring module that extracts highlight cues from video frames through two complementary pathways: sparse sampling and dense perception. Their outputs are adaptively fused to enable robust long-sequence content understanding. For Spatial Enlargement, we construct a Topology-Aware Memory Bank that hierarchically clusters historically relevant content based on topological relationships. Instead of directly expanding memory capacity, we update the encoder features of the corresponding clusters when incorporating new videos, enabling unbounded historical association without unbounded storage growth. Extensive experiments on three widely used MVPP benchmarks demonstrate that our method consistently outperforms 11 strong baselines across mainstream metrics, achieving robust improvements in both prediction accuracy and ranking consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。