arXiv:2412.08975cs.CV2024-12被引 9

用大模型生成参考图,提升视频修复的细节与精度。

Elevating Flow-Guided Video Inpainting with Reference Generation

  • 结合大模型生成参考图像,辅助缺失区域内容合成。
  • 提出单次像素拉取方法,避免误差累积并保持亚像素精度。
  • 支持2K以上高分辨率视频修复,适合真实场景应用。

视频修复(VI)是一项挑战性任务,需在帧间有效传播可见内容的同时生成原视频中不存在的新内容。本文提出一种鲁棒且实用的VI框架,利用大模型生成参考图,并结合先进的像素传播算法。通过强大的生成模型,该方法不仅显著提升了物体移除后的帧级质量,还能根据用户提供的文本提示,在缺失区域合成新内容。在像素传播方面,提出一种单次像素拉取方法,有效避免重复采样带来的误差累积,同时保持亚像素级精度。为评估真实场景下的各类VI方法,我们还构建了高质量基准HQVI,其采用透明度蒙版合成生成精心设计的视频。在公开基准和HQVI数据集上,本方法在视觉质量和指标得分上均显著优于现有方案。此外,该方法可轻松处理超过2K分辨率的视频,凸显其在真实应用中的优势。

原文摘要 · Abstract (English)

Video inpainting (VI) is a challenging task that requires effective propagation of observable content across frames while simultaneously generating new content not present in the original video. In this study, we propose a robust and practical VI framework that leverages a large generative model for reference generation in combination with an advanced pixel propagation algorithm. Powered by a strong generative model, our method not only significantly enhances frame-level quality for object removal but also synthesizes new content in the missing areas based on user-provided text prompts. For pixel propagation, we introduce a one-shot pixel pulling method that effectively avoids error accumulation from repeated sampling while maintaining sub-pixel precision. To evaluate various VI methods in realistic scenarios, we also propose a high-quality VI benchmark, HQVI, comprising carefully generated videos using alpha matte composition. On public benchmarks and the HQVI dataset, our method demonstrates significantly higher visual quality and metric scores compared to existing solutions. Furthermore, it can process high-resolution videos exceeding 2K resolution with ease, underscoring its superiority for real-world applications.

视频修复生成模型像素传播高分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。