arXiv:2507.21858cs.CV2025-07被引 1

轻量级测试时自适应,提升视频编辑一致性与抗提示过拟合能力

Low-Cost Test-Time Adaptation for Robust Video Editing

  • 通过自监督任务在推理时动态优化每段视频
  • 显著改善视频时序一致性,减少对简单提示的过拟合
  • 计算开销极低,可直接接入现有编辑模型

视频编辑是内容创作的核心环节,将原始素材转化为符合特定视觉与叙事目标的连贯作品。现有方法面临两大挑战:难以捕捉复杂运动模式导致的时间不一致,以及受限于UNet主干网络而对简单提示产生过拟合。尽管学习方法能提升编辑质量,但通常需要大量计算资源,且受限于高质量标注数据稀缺。本文提出Vid-TTA,一种轻量级测试时自适应框架,在推理阶段通过自监督辅助任务为每段测试视频个性化优化。该方法引入运动感知帧重建机制,识别并保留关键运动区域;结合提示扰动与重构策略,增强模型对多样化文本描述的鲁棒性。上述创新由元学习驱动的动态损失平衡机制协同调度,根据视频特征自适应调整优化过程。大量实验表明,Vid-TTA显著提升视频时序一致性,缓解提示过拟合,同时保持极低计算开销,可作为即插即用模块提升现有视频编辑模型性能。

原文摘要 · Abstract (English)

Video editing is a critical component of content creation that transforms raw footage into coherent works aligned with specific visual and narrative objectives. Existing approaches face two major challenges: temporal inconsistencies due to failure in capturing complex motion patterns, and overfitting to simple prompts arising from limitations in UNet backbone architectures. While learning-based methods can enhance editing quality, they typically demand substantial computational resources and are constrained by the scarcity of high-quality annotated data. In this paper, we present Vid-TTA, a lightweight test-time adaptation framework that personalizes optimization for each test video during inference through self-supervised auxiliary tasks. Our approach incorporates a motion-aware frame reconstruction mechanism that identifies and preserves crucial movement regions, alongside a prompt perturbation and reconstruction strategy that strengthens model robustness to diverse textual descriptions. These innovations are orchestrated by a meta-learning driven dynamic loss balancing mechanism that adaptively adjusts the optimization process based on video characteristics. Extensive experiments demonstrate that Vid-TTA significantly improves video temporal consistency and mitigates prompt overfitting while maintaining low computational overhead, offering a plug-and-play performance boost for existing video editing models.

视频编辑测试时适应轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。