arXiv:2508.18859cs.CV2025-08TPAMI被引 2

用快速自适应让视频稳帧生成更准更可控。

Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization

  • 测试时用低层视觉线索快速调整模型,适配每段视频
  • 单次适配就显著提升稳定性和画质,高抖动段重点优化
  • 既保全帧输出,又让用户像传统方法一样自由控制

视频稳定仍是计算机视觉中的基础问题,尤其针对全帧像素级合成的方案,其在增强稳定性的同时增加了任务复杂性。由于视频序列中运动模式和视觉内容差异大,固定参数难以实现鲁棒泛化。为此,我们提出一种新方法,在测试时快速适应每个输入视频,利用推理时的低层视觉线索提升输出的稳定性和视觉质量。显著的是,仅一次适应便带来明显性能提升。我们进一步设计了抖动定位模块与定向适应策略,聚焦高抖动段以减少适配步数并最大化稳定性。该方法使现代稳定器超越现有最先进水平,同时保持全帧特性,并提供类似经典方法的可控机制。在多样真实数据集上的大量实验表明,该方法能持续提升多种全帧合成模型的性能,涵盖定性与定量指标,包括下游应用效果。

原文摘要 · Abstract (English)

Video stabilization remains a fundamental problem in computer vision, particularly pixel-level synthesis solutions for video stabilization, which synthesize full-frame outputs, add to the complexity of this task. These methods aim to enhance stability while synthesizing full-frame videos, but the inherent diversity in motion profiles and visual content present in each video sequence makes robust generalization with fixed parameters difficult. To address this, we present a novel method that improves pixel-level synthesis video stabilization methods by rapidly adapting models to each input video at test time. The proposed approach takes advantage of low-level visual cues available during inference to improve both the stability and visual quality of the output. Notably, the proposed rapid adaptation achieves significant performance gains even with a single adaptation pass. We further propose a jerk localization module and a targeted adaptation strategy, which focuses the adaptation on high-jerk segments for maximizing stability with fewer adaptation steps. The proposed methodology enables modern stabilizers to overcome the longstanding SOTA approaches while maintaining the full frame nature of the modern methods, while offering users with control mechanisms akin to classical approaches. Extensive experiments on diverse real-world datasets demonstrate the versatility of the proposed method. Our approach consistently improves the performance of various full-frame synthesis models in both qualitative and quantitative terms, including results on downstream applications.

视频稳定元学习可控生成全帧合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。