arXiv:2603.05851cs.CV2026-03

用3D重建与扩散模型结合,实现无裁剪的稳定视频生成。

VS3R: Robust Full-frame Video Stabilization via Deep 3D Reconstruction

  • 通过前馈3D重建与生成式扩散模型协同工作
  • 在极端运动下仍保持全帧稳定,视觉质量显著提升
  • 适合需要高质量视频稳定的应用场景

视频稳定旨在减轻摄像机抖动,但面临几何鲁棒性与全帧一致性的根本矛盾。2D方法常导致剧烈裁剪,而3D技术则易受脆弱优化流程影响,在极端运动下失效。新型视图合成模型存在结构伪影和尺度感知盲区问题。为此,我们提出VS3R框架,融合前馈3D重建与生成式视频扩散模型。该流程联合估计相机参数、深度图与掩码,确保全场景可靠性,并引入混合稳定渲染(HSR)模块,融合语义与几何线索,初步缓解由姿态变换引起的视差遮挡,同时保持动态与静态一致性。最后,视频稳定驱动扩散模型(VSDM)利用上下文信息恢复裸露区域,联合优化纹理与时间一致性。整体上,VS3R在多种摄像机模型下实现了高保真全帧稳定,显著优于当前最优方法,在鲁棒性与视觉质量上均表现突出。

原文摘要 · Abstract (English)

Video stabilization aims to mitigate camera shake but faces a fundamental trade-off between geometric robustness and full-frame consistency. While 2D methods suffer from aggressive cropping, 3D techniques are often undermined by fragile optimization pipelines that fail under extreme motions. Novel view synthesis models suffer from structural artifacts and scale blindness. To bridge this gap, we propose VS3R, a framework that synergizes feed-forward 3D reconstruction with generative video diffusion. Our pipeline jointly estimates camera parameters, depth, and masks to ensure all-scenario reliability, and introduces a Hybrid Stabilized Rendering (HSR) module that fuses semantic and geometric cues to preliminarily address parallax occlusions caused by pose transformations while maintaining dynamic-static consistency. Finally, a Video Stabilization-Driven Diffusion Model (VSDM) leverages contextual information to restore disoccluded regions, jointly optimizing texture and temporal consistency. Collectively, VS3R achieves high-fidelity, full-frame stabilization across diverse camera models and significantly outperforms state-of-the-art methods in robustness and visual quality.

视频稳定3D重建扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。