arXiv:2510.04236cs.CV2025-10被引 8

Kaleido用序列到序列框架统一生成3D与视频,无需显式3D表示。

Scaling Sequence-to-Sequence Generative Neural Rendering

  • 将3D视为视频的特例,用序列到序列图像合成建模。
  • 零样本下少视角生成质量超越现有方法,多视角媲美优化类方法。
  • 基于大规模视频预训练,减少对标注3D数据依赖。

我们提出Kaleido,一组用于逼真、统一的物体级与场景级神经渲染的生成模型。Kaleido基于3D可视为视频特殊子域的原理,将3D渲染纯粹表达为序列到序列图像合成任务。通过对序列到序列生成神经渲染的系统性扩展研究,我们引入关键架构创新,使模型能够:一)在不使用显式3D表示的情况下进行生成式视角合成;二)通过掩码自回归框架,根据任意数量参考视图生成任意数量6-自由度目标视图;三)在单一解码器仅的修正流变换器中无缝统一3D与视频建模。在此统一框架下,Kaleido利用大规模视频数据进行预训练,显著提升空间一致性并降低对稀缺相机标注3D数据集的依赖——且无需任何架构修改。Kaleido在多个视角合成基准上达到新最优水平。其零样本性能在少视角设置下显著优于其他生成方法,并首次在多视角设置下达到逐场景优化方法的质量。

原文摘要 · Abstract (English)

We present Kaleido, a family of generative models designed for photorealistic, unified object- and scene-level neural rendering. Kaleido operates on the principle that 3D can be regarded as a specialised sub-domain of video, expressed purely as a sequence-to-sequence image synthesis task. Through a systemic study of scaling sequence-to-sequence generative neural rendering, we introduce key architectural innovations that enable our model to: i) perform generative view synthesis without explicit 3D representations; ii) generate any number of 6-DoF target views conditioned on any number of reference views via a masked autoregressive framework; and iii) seamlessly unify 3D and video modelling within a single decoder-only rectified flow transformer. Within this unified framework, Kaleido leverages large-scale video data for pre-training, which significantly improves spatial consistency and reduces reliance on scarce, camera-labelled 3D datasets -- all without any architectural modifications. Kaleido sets a new state-of-the-art on a range of view synthesis benchmarks. Its zero-shot performance substantially outperforms other generative methods in few-view settings, and, for the first time, matches the quality of per-scene optimisation methods in many-view settings.

神经渲染序列生成统一建模视频预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。