用分形纹理增强相机控制,让单视频重拍更真实稳定
SierpinskiCam: Camera-Controlled Video Retaking with Sierpinski Triangle Pattern Cues

- 用谢尔宾斯基三角形纹理提供强跟踪特征,适应大幅视角变化
- 通过参考视频帧绑定外观,实现无需修改模型的精准视觉一致性
- 适合影视特效与3D内容创作,尤其在复杂运动轨迹下表现优异
从单个单目视频生成用户定义相机路径下的新画面(即视频重拍),是内容创作与视觉特效中的重要但困难的问题。现有基于几何的方法通过重建4D表示并沿目标轨迹渲染来指导视频扩散模型,但当相机偏离原始轨迹时,引导效果下降,导致新暴露区域稀疏或缺失。本文提出SierpinskiCam,通过引入包含丰富可追踪特征的谢尔宾斯基穹顶纹理,增强几何引导,即使在大视角变化下仍保持稳定性。进一步设计参考视频条件机制,将源视频帧令牌附加至目标令牌序列,并使用负RoPE索引分离两流,实现外观对齐而无需架构修改或每视频适配。大量实验表明,SierpinskiCam在多种复杂重拍场景中显著提升相机可控性、几何一致性与视频质量。
原文摘要 · Abstract (English)
Generating novel renderings of a scene along user-defined camera trajectories from a single monocular video, dubbed video retaking, is a compelling but difficult problem in content creation and visual effects. Existing geometry-guided approaches reconstruct a 4D representation from the source video and render it along the target trajectory to condition video diffusion models. However, this guidance degrades as the target camera departs from the source trajectory, leaving newly revealed regions sparse or entirely missing. We propose SierpinskiCam, which addresses this limitation by augmenting geometry-based guidance with Sierpinski dome texture cues that contains rich trackable features even under large viewpoint changes. We further introduce a reference video conditioning mechanism that appends source-video tokens to the target-token sequence and separates the two streams with negative RoPE indices, enabling appearance grounding without architectural modification or per-video adaptation. Extensive experiments show that SierpinskiCam achieves significant gains in camera controllability, geometric consistency, and video quality across diverse and challenging retaking scenarios. Project page: https://hyelinnam.github.io/SierpinskiCam/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。