arXiv:2507.04293cs.ROcs.CV2025-07被引 2

AutoLayout通过双系统协作生成更真实合理的布局,解决物体漂浮重叠问题。

AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning

  • 慢系统用推理-反思-生成流程分析物体属性与空间约束
  • 快系统生成坐标与拓扑关系,经自验证闭环修正错误
  • 引入大模型驱动的动态关系库,提升布局合理性,适合机器人环境构建

自动化布局生成对具身智能与自主系统至关重要,广泛应用于虚拟环境构建到家庭机器人部署。现有方法常出现空间幻觉,难以兼顾语义准确性和物理合理性,导致物体漂浮、重叠或堆叠错位。本文提出AutoLayout,一种全自动化方法,基于双系统框架集成闭环自验证机制。慢系统采用推理-反思-生成(RRG)流程,细致提取物体属性与空间约束;快系统生成离散坐标集与拓扑关系集,并联合验证。为克服手工规则局限,进一步引入基于大模型的自适应关系库(ARL),实现布局生成与评估。通过慢-快协同推理,AutoLayout在充分权衡后高效生成布局,有效缓解空间幻觉。其自验证机制形成闭环,迭代修正潜在错误,在8个不同场景中,相比最先进方法在物理合理性、语义一致性与功能完整性上提升10.1%。

原文摘要 · Abstract (English)

The automated generation of layouts is vital for embodied intelligence and autonomous systems, supporting applications from virtual environment construction to home robot deployment. Current approaches, however, suffer from spatial hallucination and struggle with balancing semantic fidelity and physical plausibility, often producing layouts with deficits such as floating or overlapping objects and misaligned stacking relation. In this paper, we propose AutoLayout, a fully automated method that integrates a closed-loop self-validation process within a dual-system framework. Specifically, a slow system harnesses detailed reasoning with a Reasoning-Reflection-Generation (RRG) pipeline to extract object attributes and spatial constraints. Then, a fast system generates discrete coordinate sets and a topological relation set that are jointly validated. To mitigate the limitations of handcrafted rules, we further introduce an LLM-based Adaptive Relation Library (ARL) for generating and evaluating layouts. Through the implementation of Slow-Fast Collaborative Reasoning, the AutoLayout efficiently generates layouts after thorough deliberation, effectively mitigating spatial hallucination. Its self-validation mechanism establishes a closed-loop process that iteratively corrects potential errors, achieving a balance between physical stability and semantic consistency. The effectiveness of AutoLayout was validated across 8 distinct scenarios, where it demonstrated a significant 10.1% improvement over SOTA methods in terms of physical plausibility, semantic consistency, and functional completeness.

布局生成具身智能双系统自验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。