arXiv:2601.05810cs.CVcs.AI2026-01被引 4

用自然语言生成可交互的3D公寓世界,支持机器人训练

SceneFoundry: Generating Interactive Infinite 3D Worlds

  • 通过语言指令控制布局,用扩散模型填充带活动部件的家具
  • 生成场景满足物理可用性,保证行走空间与无碰撞
  • 适合做机器人学习和具身智能研究的无限3D环境

自动生成大规模、可交互且物理真实的3D环境对于推进机器人学习和具身智能至关重要。然而,现有生成方法难以捕捉真实室内空间的功能复杂性,尤其是包含可移动部件的铰接式物体。本文提出SceneFoundry,一种基于语言引导的扩散框架,能够生成具有功能性的铰接家具和语义多样的布局的公寓级3D世界,用于机器人训练。通过自然语言提示,大语言模型模块控制地板布局生成,而基于扩散的后验采样则从大规模3D资源库中高效填充带关节的资产。为确保物理可用性,SceneFoundry采用可微分引导函数调节物体数量、避免关节碰撞,并保持足够的可行走空间以支持机器人导航。大量实验表明,该框架在多种场景类型和条件下均能生成结构合理、语义连贯且功能可交互的环境,支持可扩展的具身人工智能研究。

原文摘要 · Abstract (English)

The ability to automatically generate large-scale, interactive, and physically realistic 3D environments is crucial for advancing robotic learning and embodied intelligence. However, existing generative approaches often fail to capture the functional complexity of real-world interiors, particularly those containing articulated objects with movable parts essential for manipulation and navigation. This paper presents SceneFoundry, a language-guided diffusion framework that generates apartment-scale 3D worlds with functionally articulated furniture and semantically diverse layouts for robotic training. From natural language prompts, an LLM module controls floor layout generation, while diffusion-based posterior sampling efficiently populates the scene with articulated assets from large-scale 3D repositories. To ensure physical usability, SceneFoundry employs differentiable guidance functions to regulate object quantity, prevent articulation collisions, and maintain sufficient walkable space for robotic navigation. Extensive experiments demonstrate that our framework generates structurally valid, semantically coherent, and functionally interactive environments across diverse scene types and conditions, enabling scalable embodied AI research. project page: https://anc891203.github.io/SceneFoundry-Demo/

3D生成机器人训练扩散模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。