arXiv:2509.22281cs.CVcs.RO2025-09NeurIPS被引 19

用空间推理生成符合任务的逼真桌面场景,提升机器人训练效果。

MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning

  • 通过空间推理链分解生成过程:物体推断→空间关系分析→场景图构建
  • 在10,700个合成场景上验证,生成结果与任务描述匹配度显著提升
  • 适合研究机器人交互、具身智能和场景生成的开发者与研究人员

机器人理解人类指令并执行操作任务,依赖于与任务相关的桌面场景进行训练。然而,传统方法依赖耗时的手动布局设计或纯随机布局,存在合理性不足或与任务不匹配的问题。本文提出一项新任务——面向任务的桌面场景生成,并引入包含约10,700个合成场景的MesaTask-10K数据集,所有布局由人工精心设计,确保真实感和复杂的物间关系。为弥合任务指令与场景之间的鸿沟,提出空间推理链,将生成过程分解为物体推断、空间关系推理和场景图构建三个阶段。基于此,构建了基于大语言模型的MesaTask框架,并采用DPO算法优化,生成物理上合理且与任务描述高度一致的3D桌面场景。大量实验证明,MesaTask在生成符合任务要求的逼真场景方面优于基线方法。

原文摘要 · Abstract (English)

The ability of robots to interpret human instructions and execute manipulation tasks necessitates the availability of task-relevant tabletop scenes for training. However, traditional methods for creating these scenes rely on time-consuming manual layout design or purely randomized layouts, which are limited in terms of plausibility or alignment with the tasks. In this paper, we formulate a novel task, namely task-oriented tabletop scene generation, which poses significant challenges due to the substantial gap between high-level task instructions and the tabletop scenes. To support research on such a challenging task, we introduce MesaTask-10K, a large-scale dataset comprising approximately 10,700 synthetic tabletop scenes with manually crafted layouts that ensure realistic layouts and intricate inter-object relations. To bridge the gap between tasks and scenes, we propose a Spatial Reasoning Chain that decomposes the generation process into object inference, spatial interrelation reasoning, and scene graph construction for the final 3D layout. We present MesaTask, an LLM-based framework that utilizes this reasoning chain and is further enhanced with DPO algorithms to generate physically plausible tabletop scenes that align well with given task descriptions. Exhaustive experiments demonstrate the superior performance of MesaTask compared to baselines in generating task-conforming tabletop scenes with realistic layouts. Project page is at https://mesatask.github.io/

场景生成空间推理机器人训练3D布局

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。