arXiv:2504.03875cs.CV2025-04被引 5

用随机序列生成提升3D场景理解一致性。

3D Scene Understanding Through Local Random Access Sequence Modeling

  • 局部块量化+随机序列生成,提升3D重建一致性。
  • 在新视角合成与3D物体操控上达当前最优效果。
  • 可自然扩展至自监督深度估计,适合3D视觉研究者。

从单张图像进行3D场景理解是计算机视觉中的关键问题,广泛应用于图形学、增强现实和机器人领域。尽管基于扩散模型的方法已展现潜力,但在复杂真实场景中常难以保持物体与场景的一致性。为此,我们提出一种自回归生成方法——局部随机访问序列(LRAS)建模,通过局部块量化与随机顺序序列生成实现。利用光流作为3D场景编辑的中间表示,实验表明,LRAS在新视角合成和3D物体操控任务上达到当前最优表现。此外,通过简单修改序列设计,该框架可自然延伸至自监督深度估计。在多个3D场景理解任务中均取得优异性能,为下一代3D视觉模型构建提供统一高效框架。

原文摘要 · Abstract (English)

3D scene understanding from single images is a pivotal problem in computer vision with numerous downstream applications in graphics, augmented reality, and robotics. While diffusion-based modeling approaches have shown promise, they often struggle to maintain object and scene consistency, especially in complex real-world scenarios. To address these limitations, we propose an autoregressive generative approach called Local Random Access Sequence (LRAS) modeling, which uses local patch quantization and randomly ordered sequence generation. By utilizing optical flow as an intermediate representation for 3D scene editing, our experiments demonstrate that LRAS achieves state-of-the-art novel view synthesis and 3D object manipulation capabilities. Furthermore, we show that our framework naturally extends to self-supervised depth estimation through a simple modification of the sequence design. By achieving strong performance on multiple 3D scene understanding tasks, LRAS provides a unified and effective framework for building the next generation of 3D vision models.

3D理解生成模型自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。