arXiv:2606.27718cs.CV2026-06中稿 · ECCV

提出动态轨迹扫描方法,提升复杂运动下的视频插帧精度。

MASS: Motion-Aligned Selective Scan for Refinement in Flow-Based Video Frame Interpolation

论文配图:MASS: Motion-Aligned Selective Scan for Refinement in Flow-Based Video Frame Interpolation
图 1 · 摘自论文原文
  • 按光流轨迹动态扫描特征,替代固定网格扫描
  • 在大位移场景下优于现有方法,尤其处理复杂运动时表现突出
  • 适合需要高精度插帧的视频增强与动画生成任务

视频帧插值(VFI)在处理大范围非线性运动和复杂遮挡时仍具挑战。尽管基于光流的方法广泛应用,但常因对应关系模糊而受限。现有基于选择性状态空间模型(SSM)的方法受限于静态网格扫描,无法对齐实际运动。本文提出运动对齐选择性扫描(MASS),将特征扫描从静态空间网格重构为动态运动轨迹。MASS沿每个像素的光流引导轨迹构建特征序列,并通过SSM进行聚合。具体地,引入可学习的非线性路径积分,通过残差速度更新近似复杂曲线轨迹;同时设计速度感知的SSM,根据运动幅度动态调整采样预算与步长,使快速运动区域获得更密集采样,静态区域保持高效。此外,聚合状态指导端到端的精炼模块,校正中间光流与掩码。大量实验表明,MASS在标准基准上表现优异,尤其在大位移与复杂动态场景中达到当前最优性能。

原文摘要 · Abstract (English)

Video frame interpolation (VFI) remains a challenging task, particularly when dealing with large, non-linear motions and complex occlusions. While flow-based methods are prevalent, they often struggle with ambiguous correspondences. Recent VFI methods based on selective State Space Models (SSMs) are still limited by static grid-based scanning that misaligns with physical motion. In this paper, we propose Motion-Aligned Selective Scan (MASS), a novel framework that reformulates feature scanning from static spatial grids to dynamic motion trajectories. MASS builds a feature sequence along each pixel's flow-guided trajectory and aggregates it with an SSM. Specifically, we introduce a learnable non-linear path integration to approximate complex curved trajectories via residual velocity updates, and a velocity-aware SSM that dynamically adjusts the sampling budget and step size based on motion magnitude. This adaptive strategy allocates denser sampling to fast-motion regions while keeping static regions efficient. Furthermore, the aggregated states guide a refinement module to rectify intermediate flows and masks in an end-to-end manner. Extensive experiments indicate that MASS achieves highly competitive overall performance on standard benchmarks, establishing state-of-the-art results particularly in challenging scenarios with large displacements and complex dynamics.

视频插帧光流状态空间模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。