通过分步生成与修复,实现可控且物理合理的3D室内场景合成。
ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation

- 分层检索先验布局,支持功能分组规划。
- 每步插入后轻量修复,最终全局校正,提升合理性。
- 适合需要高可控性与真实感的场景生成任务。
文本驱动的3D室内场景生成已从数据依赖的布局建模发展为基于大语言模型与视觉语言模型的开放词汇合成。然而现有方法仍受限:单次生成常产生几何无效布局,后期优化成本高且不稳定,仅靠提示的规划缺乏可复用的功能分组与物体关系先验。我们提出 extbf{ScenePilot},一种检索增强的 extbf{Grow-and-Repair}框架,将场景生成建模为先验引导的增量生长与学习修正过程。给定提示,分层检索增强规划(HRAP)模块检索房间级、组级和锚点级布局先验,支持功能分组规划。文本驱动的基线生成器按序插入物体组,同时强化多模态修复(RMR)模块在每次插入后进行轻量局部修正,并在完成后执行最终全局修复。为训练该策略,我们构建了 extbf{SceneReverse-17k},一个由17,000个高质量3D场景在位置、旋转和尺度上扰动后,通过逆操作生成可执行修复目标的修复轨迹数据集。该策略从渲染视图、场景状态、检索先验和编辑历史中预测结构化的 extit{移动--旋转--缩放}动作。结合HRAP与RMR,ScenePilot提供了一种高效替代单次生成与重型全场景优化的方法,在保持多样性的同时,显著提升物理合理性、功能连贯性与可控性。
原文摘要 · Abstract (English)
Text-driven 3D indoor scene generation has advanced from dataset-bound layout modeling to open-vocabulary synthesis with large language and vision-language models. Yet existing methods remain limited: one-pass generators often yield geometrically invalid layouts, heavy post-hoc optimization is costly and unstable, and prompt-only planners lack reusable layout priors for functional grouping and object relations. We propose \textbf{ScenePilot}, a retrieval-augmented \textbf{Grow-and-Repair} framework that formulates scene generation as prior-guided incremental growth with learned rectification. Given a prompt, the Hierarchical Retrieval-Augmented Planning (HRAP) module retrieves room-, group-, and anchor-level layout priors to support functional group planning. A text-driven base generator then inserts object groups sequentially, while the Reinforcement Multimodal Repair (RMR) module performs lightweight local correction after each insertion and a final global repair after completion. To train this policy, we construct \textbf{SceneReverse-17k}, a repair-trajectory dataset built by perturbing high-quality 3D scenes in position, rotation, and scale, then using inverse operations as executable rectification targets. The policy predicts structured \emph{move--rotate--scale} actions from rendered views, scene state, retrieved priors, and edit history. By combining HRAP with RMR, ScenePilot offers an efficient alternative to one-shot generation and heavy full-scene optimization, improving physical plausibility, functional coherence, and controllability while preserving diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。