从单张图像生成真实尺度的3D场景,解决3D数据稀缺问题
Towards Scalable Spatial Intelligence via 2D-to-3D Data Lifting
- 通过深度估计与标定联合优化,将2D图像转为带尺度的3D表示
- 生成的COCA-3D和Objects365-v2-3D数据集可支持多种3D任务
- 适合研究3D感知、多模态大模型推理与空间智能的开发者
空间智能正成为人工智能的新前沿,但受限于大规模3D数据集的缺乏。与丰富的2D图像不同,获取3D数据通常需要专用传感器和繁重的人工标注。本文提出一种可扩展的流水线,通过集成深度估计、相机标定与尺度标定,将单视图图像转化为包含点云、相机位姿、深度图和伪-RGBD的完整、真实尺度且外观逼真的3D表征。该方法弥合了海量图像资源与日益增长的空间场景理解需求之间的鸿沟。通过自动从图像生成真实尺度的3D数据,显著降低数据采集成本,为推进空间智能开辟新路径。我们发布了两个自动生成的空间数据集:COCO-3D 和 Objects365-v2-3D。大量实验证明,生成数据可有效提升从基础感知到多模态大模型推理等各类3D任务的表现。结果验证了该流水线在构建具备物理环境感知、理解与交互能力的AI系统中的有效性。
原文摘要 · Abstract (English)
Spatial intelligence is emerging as a transformative frontier in AI, yet it remains constrained by the scarcity of large-scale 3D datasets. Unlike the abundant 2D imagery, acquiring 3D data typically requires specialized sensors and laborious annotation. In this work, we present a scalable pipeline that converts single-view images into comprehensive, scale- and appearance-realistic 3D representations - including point clouds, camera poses, depth maps, and pseudo-RGBD - via integrated depth estimation, camera calibration, and scale calibration. Our method bridges the gap between the vast repository of imagery and the increasing demand for spatial scene understanding. By automatically generating authentic, scale-aware 3D data from images, we significantly reduce data collection costs and open new avenues for advancing spatial intelligence. We release two generated spatial datasets, i.e., COCO-3D and Objects365-v2-3D, and demonstrate through extensive experiments that our generated data can benefit various 3D tasks, ranging from fundamental perception to MLLM-based reasoning. These results validate our pipeline as an effective solution for developing AI systems capable of perceiving, understanding, and interacting with physical environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。