arXiv:2609.03657cs.CV2026-09

无需优化即可提升3D视频生成的几何一致性,解决稀疏视角下的重建瑕疵。

Rethinking 3D Noise: Learning 3D-Aware Video Priors via Optimization-Free Morphological Perturbations

论文配图:Rethinking 3D Noise: Learning 3D-Aware Video Priors via Optimization-Free Morphological Perturbations
图 1 · 摘自论文原文
  • 用3D高斯的尺度、旋转和删减做形态扰动,替代传统优化流程。
  • 在轻量级视频扩散模型上,深度误差降低12.5%,下游机器人任务成功率最高提升8.0%。
  • 适用于希望提升3D几何先验的视频生成与机器人控制研究者。

NeRF和3D高斯溅射(3DGS)在稀疏视角下存在严重伪影。现有生成式3D修复方法依赖成对的损坏与清晰渲染图像,需针对每个场景进行代价高昂的重建。尽管2D图像增强可即时正则化,但尚无显式的3D表示正则化手段来保持跨视角的空间一致性,而这是3D感知训练的关键。本文提出一种免优化的3D形态扰动正则化方法,通过显式3DGS将每个高斯视为基本单元(类比2D像素),在其形态参数空间中施加尺度、旋转和剪枝扰动。该方法省去了数据集构建中的逐场景3DGS优化循环,使模型在轻量级视频扩散沙盒中的诊断消融实验中学习到强于稀疏视角基线的几何先验。扩展至140亿参数视频模型(通过ControlNet),本方法在保持视觉保真度的同时,相较最先进的图像到图像3D伪影修复器,平均深度误差降低12.5%,最终在四个操作任务中的三个上,提升下游机器人策略成功率高达8.0%。

原文摘要 · Abstract (English)

3D scene representations like NeRF and 3D Gaussian Splatting (3DGS) suffer severe artifacts in sparse-view settings. Recent generative 3D artifact fixers attempt to address this, but rely on paired corrupted and clean renders requiring costly, per-scene reconstructions across varying view configurations. While 2D image augmentations act as instant regularizers, no explicit equivalents exist for 3D representations to preserve spatial consistency across views, an essential property for 3D-aware training. We propose 3D Morphological Perturbations as an optimization-free regularizer that preserves spatial consistency. Leveraging explicit 3DGS, we treat each Gaussian as a fundamental building block - analogous to a 2D pixel - and apply perturbations across its morphological parameter space via scale, rotation, and pruning. Our method eliminates per-scene 3DGS optimization loops from dataset curation while enabling models to learn stronger geometric priors than sparse-view baselines in diagnostic ablations conducted on a lightweight video diffusion sandbox. Scaled to a 14B-parameter video model via ControlNet, our approach maintains visual fidelity while reducing mean depth error by 12.5% over state-of-the-art image-to-image 3D artifact refiners, ultimately boosting downstream robotics policy success rates by up to 8.0% across 3 of 4 manipulation tasks.

3D生成视频扩散机器人控制高斯溅射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。