arXiv:2608.18840cs.ROcs.CV2026-08

用代码生成可交互的3D场景,让物体使用更智能。

Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction

论文配图:Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction
图 1 · 摘自论文原文
  • 以任务为中心推理物体功能,动态生成可用对象
  • 通过规则化交互实现物体状态更新与因果关联
  • 解决物体朝向模糊问题,适合机器人训练与仿真

室内场景合成对具身智能、机器人操作和基于仿真的策略学习至关重要。现有基于代码的场景生成方法主要关注视觉构建与物体层级的运动,却忽视了场景的功能性使用。为此,我们提出RoomWright——一种完全以代码表示、面向具身交互的使用驱动式3D场景生成框架。RoomWright通过任务中心化的物体推理机制,将每个锚点视为任务核心,自动引入任务所需对象及其可用性;同时,代码代理将多步交互编译为触发-条件-效果规则,更新结构化物体状态,捕捉物体间的因果依赖。此外,针对操作物朝向难以从像素恢复的问题,采用注释引导的使用驱动方式优化朝向设定。大量实验表明,生成的场景具备可执行性、可编辑性和仿真就绪性,为具身智能与策略学习提供可交互环境。

原文摘要 · Abstract (English)

Indoor scene synthesis provides essential environments for embodied AI, robotic manipulation, and simulation-based policy learning. Recent code-based scene generation methods produce editable and extensible environments, yet they remain focused on visual construction and object-level articulation, leaving the functional usage of scenes largely unmodeled. To address this problem, we present RoomWright, an agentic usage-driven framework for generating 3D scenes represented entirely as code for embodied interaction. RoomWright performs usage-driven object reasoning, which treats each anchor as a task centre and admits task-required objects and their affordances. A code agent further enables multi-part interaction by compiling each interaction into a trigger, condition, effect rule that updates structured object states, capturing causal dependencies across objects. Moreover, since manipuland orientation is ambiguous and hard to recover from pixels, RoomWright alleviates this via annotation-informed usage-guided orientation. Extensive experiments demonstrate the effectiveness of our method. The resulting scenes are executable, editable, and simulation-ready, providing interactive environments for embodied AI and policy learning.

具身智能场景生成代码化建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。