用少量无序令牌压缩3D场景,支持高效生成与新视角渲染。
SceneTok: A Compressed, Diffusable Token Space for 3D Scenes
- 将多视角场景编码为可置换的紧凑无结构令牌,脱离空间网格。
- 压缩率比现有方法高1-3个数量级,重建质量达顶尖水平。
- 5秒内完成高质量场景生成,适合快速原型与交互应用。
我们提出SceneTok,一种新型分词器,将场景的视图集合编码为压缩且可扩散的无结构令牌集。现有3D场景表示与生成方法多依赖3D数据结构或视图对齐场,而我们的方法首次将场景信息编码为少量排列不变的令牌,与空间网格解耦。该分词器基于多个上下文视图预测场景令牌,并通过轻量级修正流解码器渲染新视角。实验表明,该表示压缩率高达1-3个数量级,同时保持最先进重建质量;支持从偏离输入轨迹的新路径渲染,且解码器能优雅处理不确定性。此外,高度压缩的无结构隐式令牌集可在5秒内实现简单高效的场景生成,显著优于以往范式在质量和速度上的权衡。
原文摘要 · Abstract (English)
We present SceneTok, a novel tokenizer for encoding view sets of scenes into a compressed and diffusable set of unstructured tokens. Existing approaches for 3D scene representation and generation commonly use 3D data structures or view-aligned fields. In contrast, we introduce the first method that encodes scene information into a small set of permutation-invariant tokens that is disentangled from the spatial grid. The scene tokens are predicted by a multi-view tokenizer given many context views and rendered into novel views by employing a light-weight rectified flow decoder. We show that the compression is 1-3 orders of magnitude stronger than for other representations while still reaching state-of-the-art reconstruction quality. Further, our representation can be rendered from novel trajectories, including ones deviating from the input trajectory, and we show that the decoder gracefully handles uncertainty. Finally, the highly-compressed set of unstructured latent scene tokens enables simple and efficient scene generation in 5 seconds, achieving a much better quality-speed trade-off than previous paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。