通过多轮强化学习实现3D室内场景的自我修正生成,减少空间错乱。
SceneReVis: A Self-Reflective Vision-Grounded Framework for 3D Indoor Scene Synthesis via Multi-turn RL
- 采用诊断-行动循环,逐步检查并修正空间冲突。
- 在SceneChain-12k数据集上训练,生成高质量且无碰撞的场景。
- 适合需要高精度3D场景生成的研究与应用者。
现有的一次性3D场景生成方法常因缺乏反思性推理而产生空间幻觉(如物体碰撞)。为此,我们提出SceneReVis,一种基于视觉引导的自反思框架,通过多轮强化学习的“诊断-行动”循环,利用多模态反馈显式检测并解决空间冲突。为支持这一渐进式范式,我们构建了SceneChain-12k,一个通过新型逆向工程流程生成的大规模因果建造轨迹数据集。我们进一步提出两阶段训练方案:从监督微调过渡到代理式强化学习,使模型演化为积极的空间规划者。大量实验表明,SceneReVis在高保真生成和目标导向优化方面达到当前最优性能,并在长尾领域表现出强泛化能力。
原文摘要 · Abstract (English)
Current one-pass 3D scene synthesis methods often suffer from spatial hallucinations, such as collisions, due to a lack of deliberative reasoning. To bridge this gap, we introduce SceneReVis, a vision-grounded self-reflection framework that employs an iterative ``diagnose-and-act'' loop to explicitly intercept and resolve spatial conflicts using multi-modal feedback. To support this step-wise paradigm, we construct SceneChain-12k, a large-scale dataset of causal construction trajectories derived through a novel reverse engineering pipeline. We further propose a two-stage training recipe that transitions from Supervised Fine-Tuning to Agentic Reinforcement Learning, evolving the model into an active spatial planner. Extensive experiments demonstrate that SceneReVis achieves state-of-the-art performance in high-fidelity generation and goal-oriented optimization, with robust generalization to long-tail domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。