MUSE让3D场景编辑像拼乐高,能记住需求、自动纠错。
MUSE: Agentic 3D Scene Authoring via Memory-Grounded Incremental Requirement Satisfaction

- 用多智能体分步构建场景,记忆跟踪每一步要求
- 编辑时99.9%内容不变,错误率仅0.6%
- 适合需要精准修改的数字内容创作与仿真设计
文本驱动的3D场景生成在数字内容创作、具身AI模拟和交互设计中前景广阔,但实际工作流常需在不破坏已有内容的前提下修改或扩展场景。现有方法虽能生成逼真且结构合理的场景,但缺乏对需求级别的状态追踪,局部失败往往导致全场景重做或人工干预。为此,本文将可控3D场景创作建模为增量式需求满足过程,统一了构建与编辑。提出MUSE——一个基于记忆的多智能体框架:建筑师将指令转化为结构化需求,雕塑家执行局部操作,检查员验证每一步并更新工作记忆、场景记忆和技能记忆。为评估需求级控制力与保留感知编辑能力,引入AuthorBench,包含145个约束构建案例和1,584个保留感知编辑案例,并配有外部结构化验证。在完整构建任务上,MUSE将全目标成功率从37.9提升至80.7,表面约束满足率从35.0提升至92.6;在分层240例编辑测试中,实现49.6%全目标成功率、99.9%内容保留率和仅0.6%的意外变更率。人类评估支持其更契合用户意图,下游导航代理测试也显示更强的空间稳定性。消融实验验证了记忆设计的有效性,确立MUSE在可控3D场景创作中的有效性。
原文摘要 · Abstract (English)
Text-driven 3D scene generation is a promising technique for digital content creation, embodied AI simulation, and interactive design, yet practical workflows often require refining, extending, or correcting existing scenes while preserving non-target content. Existing methods can produce realistic and structurally plausible scenes, but they generally lack editability with requirement-level state tracking, so part-level failures often lead to full-scene regeneration or manual intervention. To tackle this challenge, we formulate controllable 3D scene authoring as incremental requirement satisfaction, unifying construction and editing. In this paper, we present MUSE, a memory-grounded multi-agent framework in which an Architect compiles instructions into structured requirements, a Sculptor executes local scene operations, and an Inspector verifies each step while updating Working, Scene, and Skill Memory. To evaluate requirement-level controllability and preservation-aware editing, we introduce AuthorBench, offering 145 constrained construction cases and a 1,584-case preservation-aware editing pool paired with external structured checks. On full construction cases, MUSE improves All-Goal success from 37.9 to 80.7 and surface-constraint fulfillment from 35.0 to 92.6 over the strongest baseline. On a stratified 240-case editing test split, MUSE achieves 49.6 All-Goal success, 99.9 preservation rate, and only 0.6 unintended change rate. Beyond automated metrics, human evaluations on compared local-editing baselines support stronger alignment with user intent, and downstream navigation-proxy tests indicate stronger spatial stability. Combined with ablations validating our memory designs, these results establish MUSE as an effective framework for controllable 3D scene authoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。