用动画网格+相机轨迹生成视频,提升4D渲染精度。
Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh

- 用神经4D G-buffer融合运动网格与相机轨迹,指导视频模型生成。
- 在68个测试案例上,峰值信噪比达25.36,比基线高1.54 dB。
- 关键创新是追踪+世界坐标,避免深度混淆运动信息。
预训练视频扩散模型在已知动画网格、相机轨迹和参考图像的情况下可充当渲染器。这种4D生成渲染场景引发核心问题:何种图像条件能让视频主干同时遵循相机运动与场景内部动画?我们提出DAR,一种参考引导的渲染器,将Wan2.2相机控制从仅基于普吕克射线扩展为相机与几何联合接口。DAR从动画网格投影出神经4D G-buffer(包含追踪、世界坐标和法向),并通过加宽的控制适配器注入,同时保留预训练图像到视频先验。核心设计是追踪与世界坐标:追踪识别应携带外观的持续表面元素,世界坐标给出其当前场景位置,法向提供局部形状。理论上,深度加校准射线可恢复3D,但深度是依赖相机的图表,混杂了相机与物体运动。在68项测试的DAR-4D基准上,LoRA微调的DAR达到PSNR 23.22,SSIM 0.895,LPIPS 0.134;全量微调达PSNR 25.36,SSIM 0.917,较现成的Wan2.2-Depth提升1.54 dB PSNR。对比实验表明,以深度替代世界坐标导致每个检查点的PSNR下降1.26–1.55 dB,证实追踪+世界坐标的4D渲染有效性。
原文摘要 · Abstract (English)
Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a reference image. This 4D generative rendering setting raises a representation question: what image-format condition lets a video backbone obey both camera motion and scene-internal animation? We propose DAR, a reference-guided renderer that extends Wan2.2 camera control from Plücker rays alone to a joint camera-plus-geometry interface. DAR projects a neural 4D G-buffer (tracking, world position, and normal) from the animated mesh and injects it through a widened control adapter while preserving the pretrained image-to-video prior. The central design choice is the pair of tracking and world position. Tracking identifies the persistent surface element that should carry appearance; world position gives its current scene-coordinate state; normal supplies local shape. Depth plus calibrated rays can recover 3D in principle, but depth is a camera-dependent chart in which camera and object motion are mixed. On the 68-case DAR-4D benchmark, LoRA DAR reaches PSNR 23.22, SSIM 0.895, and LPIPS 0.134, improving over off-the-shelf Wan2.2-Depth by 1.54 dB PSNR; a full fine-tune reaches PSNR 25.36 and SSIM 0.917. Matched ablations show that replacing world position by depth reduces PSNR by 1.26--1.55 dB at every checkpoint, supporting tracking+world-position correspondence as a practical 4D rendering condition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。