arXiv:2605.15843cs.CV2026-05被引 1

将静态3D世界转化为可编辑互动的物体级场景

WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenes

论文配图:WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenes
图 1 · 摘自论文原文
  • 用多模态智能体分解场景,识别可操作物体
  • 重建对齐的物体网格并修复背景,支持碰撞交互
  • 适合沉浸式内容创作与具身模拟场景

基于生成式场景合成的3D世界建模系统(如Marble)能生成连贯可探索的3D环境,但其输出通常是静态的单体资产,编辑性和物理交互能力有限,难以用于沉浸式内容创作和具身模拟。为此,我们提出WorldAct框架,将静态生成的3D世界转换为可编辑、可交互的物体级场景。WorldAct通过多模态智能体引导场景分解,识别可操作物体,重建几何对齐的物体级网格以支持交互,并通过3D补全恢复残留背景。生成场景支持物体级编辑、碰撞感知操作及具身任务执行,同时保持全局场景一致性。实验表明,相比原始生成场景,WorldAct实现了更丰富的交互场景,为可编辑、可交互3D世界模型提供了可行路径。

原文摘要 · Abstract (English)

Recent 3D world modeling systems based on generative scene synthesis, such as Marble, can create coherent and explorable 3D environments, yet their outputs are typically static monolithic assets with limited editability and physical interaction. This restricts their use in immersive content creation and embodied simulation, where generated worlds must be actively modified and manipulated. To tackle this challenge, we present WorldAct, a framework that converts static generated 3D worlds into editable and interaction-ready scenes. WorldAct uses a multimodal agent to guide scene decomposition, identify actionable objects, reconstruct geometrically aligned object-level meshes for interaction, and restore the residual background via 3D inpainting. The resulting scenes support object-level editing, collision-aware manipulation, and embodied task execution while preserving global scene coherence. Experiments show that WorldAct enables richer interaction scenarios than the original generated scenes, suggesting a practical path toward editable and interactive 3D world models.

3D生成交互场景物体级建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。