arXiv:2608.29519cs.CV2026-08

让3D房间自动满足功能需求,比传统生成更智能高效。

FuncRoom-Agent: Sequential Feed-Forward 3D Functional Indoor Scene Generation

论文配图:FuncRoom-Agent: Sequential Feed-Forward 3D Functional Indoor Scene Generation
图 1 · 摘自论文原文
  • 用分层指令语言构建可执行的房间程序,明确物体间几何与功能关系。
  • 推理时无需迭代修正,单次顺序生成,速度提升显著。
  • 结合功能、几何、未来可建性反馈优化生成质量,适合场景设计应用。

我们提出功能房间生成(Function-Room Generation),一种新的室内3D场景生成范式,旨在生成支持明确功能目标的房间,而非仅视觉合理布局。现有代理和可执行方法虽增强可控性,但依赖昂贵的生成-评估-修正循环,导致效率低下。为此,本文提出三项技术贡献:首先,设计递归领域特定语言(DSL),有效组织从房间结构、大型家具到密集支撑面与嵌套小物体的层级对象组合,将房间表示为具有显式几何与功能关系的阶段性可执行程序;其次,提出序列前馈场景构建框架,将递归构建轨迹提炼为场景构建专家,在推理时分阶段生成可执行DSL代码,由确定性执行器直接实例化,无需教师代理、在线评判器或迭代修复;第三,引入基于执行的流程奖励框架ScenePRM,通过功能、几何、关系及未来可建性反馈进行强化学习优化专家。我们进一步建立面向功能的基准,结果表明在通用室内场景生成与功能房间生成上均达到最先进性能,实现更强的功能完整性、关系正确性、几何可执行性与生成效率。

原文摘要 · Abstract (English)

We introduce Function-Room Generation, a new indoor 3D scene generation setting that creates rooms supporting explicit functional goals rather than merely visually plausible layouts. Existing agentic and executable methods improve controllability, but often depend on costly test-time generate--evaluate--revise loops, making functional room generation slow and computationally expensive. We address this challenge with three technical contributions. First, we design a recursive domain-specific language to effectively organize the hierarchical object compositions required by functional rooms, from room structure and major furniture to dense support-surface and nested small objects. It represents rooms as staged executable programs with explicit geometric and functional relations. Second, we propose a sequential feed-forward scene construction framework that distills recursive construction traces into a scene construction expert. At inference time, the expert writes executable DSL code stage by stage, and a deterministic executor directly instantiates each stage without teacher agents, online critics, or iterative repair. Third, we introduce ScenePRM, an execution-grounded process reward framework that improves the expert through reinforcement learning with functional, geometric, relational, and future-constructability feedback. We further establish a function-oriented benchmark and show state-of-the-art performance on both general indoor scene generation and function-room generation, achieving stronger functional completeness, relation correctness, geometric executability, and generation efficiency.

3D生成功能驱动智能场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。