用2D视觉语言模型生成3D人类场景互动数据,解决真实数据稀缺问题。
InHabit: Leveraging Image Foundation Models for Scalable 3D Human Placement

- 通过渲染-生成-提升三步法,将2D上下文知识迁移至3D场景
- 构建78000样本的高质量3D人-场景交互数据集,覆盖800个建筑级场景
- 适用于需要真实感人形交互的机器人、虚拟现实与三维重建任务
训练具身智能体像人类一样理解三维场景,需要大规模的人类与多样化环境有意义交互的数据,但此类数据稀缺。真实采集成本高且局限于受控环境,现有合成数据集依赖简单几何规则,忽略丰富场景上下文。相比之下,互联网规模训练的2D基础模型已具备人类-环境交互的常识知识。为将此知识迁移到3D,我们提出InHabit,一种自动且可扩展的3D场景人物填充数据生成器。InHabit遵循渲染-生成-提升原则:给定渲染的3D场景,视觉语言模型提出语境相关的动作,图像编辑模型插入人物,优化过程将编辑结果提升为与场景几何对齐的物理合理SMPL-X人体。应用于Habitat-Matterport3D,InHabit生成了InHabitants,首个大规模逼真3D人-场景交互数据集,包含约800个建筑尺度场景中的78000个样本,涵盖完整3D几何、SMPL-X人体和图像。在标准训练数据中加入InHabitants后,显著提升了基于RGB的3D人体-场景重建与接触估计性能;感知用户研究显示,我们的数据在78%情况下优于先前方法。
原文摘要 · Abstract (English)
Training embodied agents to understand 3D scenes as humans do requires large-scale data of people meaningfully interacting with diverse environments, yet such data is scarce. Real-world capture is costly and limited to controlled settings, while existing synthetic datasets rely on simple geometric heuristics, ignoring rich scene context. In contrast, 2D foundation models trained at internet scale have acquired commonsense knowledge of human-environment interactions. To transfer this knowledge to 3D, we introduce InHabit, an automatic and scalable data generator for populating 3D scenes with interacting humans. InHabit follows a render-generate-lift principle: given a rendered 3D scene, a vision-language model proposes contextually meaningful actions, an image-editing model inserts a human, and an optimization procedure lifts the edited result into physically plausible SMPL-X bodies aligned with the scene geometry. Applied to Habitat-Matterport3D, InHabit produces InHabitants, the first large-scale photorealistic 3D human-scene interaction dataset, with 78K samples across $\sim$800 building-scale scenes with complete 3D geometry, SMPL-X bodies, and images. Augmenting standard training data with InHabitants improves RGB-based 3D human-scene reconstruction and contact estimation, and in a perceptual user study our data is preferred in 78% of cases over prior art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。