让3D场景生成听懂人话,自动满足通行、可达等实际需求。
iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

- 用迭代式智能体强化学习,把自然语言指令转为可执行的场景约束
- 在通行、可达、留空任务上约束满足率显著提升,生成多样性不降
- 适合需要真实可用合成数据的机器人训练与视觉模型测试
合成3D场景正被广泛用于计算机视觉和具身AI的数据生成,但现有生成器通常只优化视觉逼真度,难以可靠满足任务关键的功能性约束。这种错配限制了合成数据在下游训练中的实用性,而无障碍通行、可访问性和空间规则合规性往往是必要条件。我们提出iARCS,一种基于迭代智能体强化学习的框架,可将预训练场景生成器适配至自然语言任务需求。iARCS采用两阶段策略:首先通过通用奖励预训练提升物理合理性与布局质量;随后利用大模型生成的奖励程序进行任务特定微调,并根据训练反馈迭代优化。实验表明,iARCS在通行性、可达性及留空任务上约束满足率显著提高,实现了有效的任务定制化约束优化,且保持了竞争性的场景多样性。进一步验证显示,iARCS生成的数据能有效提升基础生成器性能,证明其不仅是可控编辑工具,更是一种实用的合成数据生成方案。
原文摘要 · Abstract (English)
Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits the usefulness of synthetic data for downstream training, where accessibility, traversability, and spatial rule compliance are often essential. We present iARCS, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to natural-language task requirements. iARCS uses a two-stage strategy: universal-reward pretraining to improve physical plausibility and layout quality, followed by task-specific fine-tuning with LLM-generated reward programs that are iteratively refined from training feedback. Experiments show improved constraint fidelity on walkability, reachability, and clearance-focused tasks, effective task-specific constraint optimization, and competitive scene diversity. We further show that data generated by iARCS improves a base generator, supporting its value as a practical synthetic data generation tool rather than only a controllable scene editing method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。