用状态空间模型实现几何感知的视频微动放大,更真实且高效。
GeoMag: Geometric-Aware Video Motion Magnification via State Space Model

- 基于状态空间模型,实现全局一致的运动放大。
- 在合成与真实数据上均优于现有方法,减少伪影并提升结构一致性。
- 构建了含复杂几何变换的大型合成数据集Geo-200K,增强训练多样性。
视频微动放大可揭示人眼难以察觉的动态,但复杂几何变换下常出现结构不一致问题。现有学习方法普遍存在卷积网络全局上下文有限与变压器计算成本高的矛盾。此外,当前训练多依赖简单线性运动,无法捕捉真实视频中的几何与成像复杂性。为此,我们提出GeoMag,一种基于状态空间模型的几何感知视频微动放大框架,实现具有线性复杂度的全局一致运动放大。我们进一步构建了包含丰富几何变换和传感器级退化的大型合成数据集Geo-200K,提升训练信号的多样性和真实性。在合成与真实世界基准上的大量实验表明,GeoMag在视觉保真度和计算效率方面持续优于先前方法,产生更少伪影并保持更好结构一致性。
原文摘要 · Abstract (English)
Video Motion Magnification (VMM) reveals imperceptible dynamics but often suffers from structural inconsistencies under complex geometric transformations. Existing learning-based methods generally face a trade-off between the limited global context of CNNs and the high computational cost of Transformers. In addition, current training protocols, largely dominated by simple linear motion, fail to capture the geometric and imaging complexities encountered in real-world videos. To address these issues, we propose GeoMag, a geometric-aware VMM framework built upon State Space Models to achieve globally consistent motion amplification with linear complexity. We further construct Geo-200K, a large-scale synthetic dataset that introduces rich geometric transformations together with sensor-realistic degradations, improving the diversity and realism of training signals. Extensive experiments on synthetic and real-world benchmarks show that GeoMag consistently outperforms prior methods in visual fidelity and computational efficiency, while producing fewer artifacts and better structural consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。