arXiv:2608.26947cs.ROcs.CV2026-08

用自然语言生成可编辑的动态4D仿真环境,支持智能体导航测试

4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied Simulation

论文配图:4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied Simulation
图 1 · 摘自论文原文
  • 输入文本/草图/照片,自动生成带物理属性的可编辑4D场景
  • 构建4DSynth-Nav基准,三难度下视觉语言模型多数任务失败
  • 环境可控可复现,适合训练与评估具身智能体

具身智能体需要视觉多样、物理交互且随时间变化的环境。过程式模拟器可生成大规模互动场景,近期4D生成器也实现了生动的视觉动态。但将这些特性整合到单一环境中仍需大量人工操作,且结果难以编辑或控制以规模化复用。我们提出4DSynth,一种可控的过程式系统,能将自然语言描述、蓝图掩码或单张照片转化为带有显式几何结构、动画角色、无碰撞轨迹和物理就绪状态的可编辑4D环境。多个场景路径共享同一几何基础表示,因此同一管道可处理动画、相机规划、渲染和任务生成。为验证完整流程,我们构建了完全由4DSynth生成的4DSynth-Nav交互式导航基准。两个视觉-语言模型在三个难度层级上均多数任务失败,且在早期子任务即停滞。相同的可控性使得每个失败可重现,每条难度轴均可独立调节。本文既提供了可控生成管道,也展示了其赋能的可扩展基准,为开发和评估具身智能体提供了实用基础。

原文摘要 · Abstract (English)

Embodied agents need environments that are visually diverse, physically interactive, and changing over time. Procedural simulators can generate large interactive scene collections, and recent 4D generators produce compelling visual dynamics. Combining these properties in one environment, however, still demands extensive manual effort, and the result is rarely editable or controllable enough to reuse at scale. We present 4DSynth, a controllable procedural system that turns a natural-language description, a blueprint mask, or a single photograph into an editable 4D environment with explicit geometry, animated actors, collision-free trajectories, and physics-ready simulation state. Multiple scene routes share one geometry-grounded representation, so the same pipeline handles animation, camera planning, rendering, and task generation. To validate the full pipeline, we construct 4DSynth-Nav, an interactive navigation benchmark generated entirely from 4DSynth's procedural scenes. Two vision-language models evaluated across three difficulty tiers both fail the majority of tasks and stall after early subtasks. The same procedural controllability that produces these environments also makes each failure reproducible and each difficulty axis independently tunable. This paper presents both a controllable generation pipeline and the scalable benchmark it enables, offering a practical foundation for developing and evaluating embodied agents.

具身智能程序化生成4D场景仿真基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。