arXiv:2603.22279cs.CVcs.AI2026-03被引 3

用结构化推理实现精准的文本控制空间编辑

3D-Layout-R1: Structured Reasoning for Language-Instructed Spatial Editing

  • 通过场景图推理实现文本引导的空间布局更新
  • 比基线方法提升15%交并比,中心位置误差降低25%
  • 适合需要精确空间关系控制的视觉编辑任务

大型语言模型和视觉语言模型虽具备出色推理能力,但在细粒度视觉编辑中仍面临空间理解与布局一致性挑战。本文提出一种结构化推理框架,通过场景图推理实现文本条件下的空间布局编辑。给定输入场景图和自然语言指令,模型在图上推理生成满足文本条件且保持空间一致性的新场景图。通过显式结构化关系引导推理过程,提升了可解释性与空间关系控制力。我们在一个包含排序、空间对齐和房间编辑任务的新基准上评估该方法,训练范式使平均交并比提升15%,中心距离误差降低25%,相比最优零样本大模型,最高实现20%的mIoU提升,显著增强空间精度。

原文摘要 · Abstract (English)

Large Language Models (LLMs) and Vision Language Models (VLMs) have shown impressive reasoning abilities, yet they struggle with spatial understanding and layout consistency when performing fine-grained visual editing. We introduce a Structured Reasoning framework that performs text-conditioned spatial layout editing via scene-graph reasoning. Given an input scene graph and a natural-language instruction, the model reasons over the graph to generate an updated scene graph that satisfies the text condition while maintaining spatial coherence. By explicitly guiding the reasoning process through structured relational representations, our approach improves both interpretability and control over spatial relationships. We evaluate our method on a new text-guided layout editing benchmark encompassing sorting, spatial alignment, and room-editing tasks. Our training paradigm yields an average 15% improvement in IoU and 25% reduction in center-distance error compared to Chain of Thought Fine-tuning (CoT-SFT) and vanilla GRPO baselines. Compared to SOTA zero-shot LLMs, our best models achieve up to 20% higher mIoU, demonstrating markedly improved spatial precision.

空间编辑结构化推理场景图多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。