arXiv:2604.21776cs.CV2026-04被引 1

用自监督方法从单视频生成多视角,实现动态场景的精准重拍。

Reshoot-Anything: A Self-Supervised Model for In-the-Wild Video Reshooting

论文配图:Reshoot-Anything: A Self-Supervised Model for In-the-Wild Video Reshooting
图 1 · 摘自论文原文
  • 从单个视频提取随机轨迹生成伪多视角三元组作为训练数据。
  • 在复杂动态场景中实现最佳时序一致性与高保真新视角合成。
  • 适合需要真实世界视频重拍的视觉生成研究者使用。

动态视频重拍中的精确相机控制受限于非刚性场景配对多视角数据的严重匮乏。我们提出一种可扩展的自监督框架,能利用互联网规模的单目视频进行训练。核心贡献是生成伪多视角训练三元组,包含源视频、几何锚点和目标视频。通过从单个输入视频中提取不同且平滑的随机行走裁剪轨迹作为源与目标视图,锚点则通过前向扭曲源视频首帧并结合稠密跟踪场合成,有效模拟推理时预期的畸变点云输入。由于独立裁剪策略引入空间错位与人工遮挡,模型无法直接复制当前源帧信息,必须主动路由并重投影源视频中不同时间与视角的缺失高保真纹理以重建目标。推理时,经过最小调整的扩散变压器利用4D点云锚点,在复杂动态场景中实现了最先进的时序一致性、鲁棒相机控制与高保真新视角合成。

原文摘要 · Abstract (English)

Precise camera control for reshooting dynamic videos is bottlenecked by the severe scarcity of paired multi-view data for non-rigid scenes. We overcome this limitation with a highly scalable self-supervised framework capable of leveraging internet-scale monocular videos. Our core contribution is the generation of pseudo multi-view training triplets, consisting of a source video, a geometric anchor, and a target video. We achieve this by extracting distinct smooth random-walk crop trajectories from a single input video to serve as the source and target views. The anchor is synthetically generated by forward-warping the first frame of the source with a dense tracking field, which effectively simulates the distorted point-cloud inputs expected at inference. Because our independent cropping strategy introduces spatial misalignment and artificial occlusions, the model cannot simply copy information from the current source frame. Instead, it is forced to implicitly learn 4D spatiotemporal structures by actively routing and re-projecting missing high-fidelity textures across distinct times and viewpoints from the source video to reconstruct the target. At inference, our minimally adapted diffusion transformer utilizes a 4D point-cloud derived anchor to achieve state-of-the-art temporal consistency, robust camera control, and high-fidelity novel view synthesis on complex dynamic scenes.

视频重拍自监督学习扩散模型4D建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。