arXiv:2512.03014cs.CV2025-12NeurIPS

让图像模型稳定生成视频,抗干扰还能保持画质。

Instant Video Models: Universal Adapters for Stabilizing Image-Based Networks

  • 插入通用适配器,无需重训练主模型即可稳定视频输出
  • 在多种图像任务中显著减少帧间闪烁,提升抗压缩/噪声/恶劣天气能力
  • 适合想快速改造现有图像模型为视频模型的研究者和开发者

将基于帧的图像模型应用于视频时,常出现时间不一致问题,如帧间闪烁。当输入含时变干扰时,问题更严重。本文提出一种通用方法,通过在任意架构中插入稳定性适配器,并使用冻结主模型的高效训练流程,实现视频推理的稳定与鲁棒。提出统一的准确-稳定-鲁棒性损失函数,理论分析其在特定条件下可生成良好稳定器训练。实验验证该方法在去噪(NAFNet)、图像增强(HDRNet)、单目深度(Depth Anything v2)和语义分割(DeepLabv3+)等任务上的有效性,显著提升时间一致性与对压缩伪影、噪声及恶劣天气的鲁棒性,同时保持或改善预测质量。

原文摘要 · Abstract (English)

When applied sequentially to video, frame-based networks often exhibit temporal inconsistency - for example, outputs that flicker between frames. This problem is amplified when the network inputs contain time-varying corruptions. In this work, we introduce a general approach for adapting frame-based models for stable and robust inference on video. We describe a class of stability adapters that can be inserted into virtually any architecture and a resource-efficient training process that can be performed with a frozen base network. We introduce a unified conceptual framework for describing temporal stability and corruption robustness, centered on a proposed accuracy-stability-robustness loss. By analyzing the theoretical properties of this loss, we identify the conditions where it produces well-behaved stabilizer training. Our experiments validate our approach on several vision tasks including denoising (NAFNet), image enhancement (HDRNet), monocular depth (Depth Anything v2), and semantic segmentation (DeepLabv3+). Our method improves temporal stability and robustness against a range of image corruptions (including compression artifacts, noise, and adverse weather), while preserving or improving the quality of predictions.

视频生成稳定性图像模型适配器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。