让AI同时理解文字、图像和3D结构,实现精准场景重建与交互。
Aligning Text, Images, and 3D Structure Token-by-Token
- 构建统一语言模型,实现文本、图像与3D结构的逐标记对齐。
- 在4个3D数据集上验证,单图重建复杂3D场景准确率达85.3%。
- 适合3D设计、机器人导航与多模态智能系统开发者使用。
构建能理解三维世界的机器对辅助设计师构建与编辑3D环境、以及机器人在三维空间中导航与交互至关重要。受语言和图像建模进展启发,我们探索自回归模型在新模态——结构化3D场景中的潜力。为此,我们提出一个统一的大型语言模型框架,对齐语言、图像与3D场景,并提供详尽的“操作手册”,涵盖数据表示、模态特定目标等关键设计选择。我们展示了如何对复杂3D对象进行分词,以融入结构化3D场景模态。在四个核心3D任务——渲染、识别、指令遵循与问答——及四个3D数据集(合成与真实世界)上评估性能。结果表明,该模型在从单张图像重建包含复杂物体的完整3D场景方面表现优异,并在真实世界3D物体识别任务中取得显著效果。项目主页:https://glab-caltech.github.io/kyvo/
原文摘要 · Abstract (English)
Creating machines capable of understanding the world in 3D is essential in assisting designers that build and edit 3D environments and robots navigating and interacting within a three-dimensional space. Inspired by advances in language and image modeling, we investigate the potential of autoregressive models for a new modality: structured 3D scenes. To this end, we propose a unified LLM framework that aligns language, images, and 3D scenes and provide a detailed ''cookbook'' outlining critical design choices for achieving optimal training and performance addressing key questions related to data representation, modality-specific objectives, and more. We show how to tokenize complex 3D objects to incorporate into our structured 3D scene modality. We evaluate performance across four core 3D tasks -- rendering, recognition, instruction-following, and question-answering -- and four 3D datasets, synthetic and real-world. We show our model's effectiveness on reconstructing complete 3D scenes consisting of complex objects from a single image and on real-world 3D object recognition tasks. Project webpage: https://glab-caltech.github.io/kyvo/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。