让大模型自动把抽象任务拆解成与3D场景匹配的具体步骤
ASHiTA: Automatic Scene-grounded HIerarchical Task Analysis
- 用大模型交替生成任务分解和场景图,实现任务与环境的动态对齐
- 在复杂任务拆解上优于纯大模型基线,接地性能媲美顶尖方法
- 适合需要精准任务规划的机器人、虚拟助手等场景
尽管近期场景重建与理解研究在将自然语言与物理3D环境对齐方面取得进展,但将抽象的高层次指令与3D场景对齐仍具挑战。高层次指令可能不显式提及场景中的语义元素,且将高层任务分解为更具体的子任务(即层次化任务分析)也依赖具体环境。本文提出ASHiTA,首个基于3D场景图生成与环境对齐的任务层次结构的框架。ASHiTA通过交替使用大模型辅助的层次化任务分析生成任务分解,以及任务驱动的3D场景图构建生成环境表征。实验表明,ASHiTA在将高层次任务分解为环境相关子任务方面显著优于大模型基线,并能实现与顶尖方法相当的接地性能。
原文摘要 · Abstract (English)
While recent work in scene reconstruction and understanding has made strides in grounding natural language to physical 3D environments, it is still challenging to ground abstract, high-level instructions to a 3D scene. High-level instructions might not explicitly invoke semantic elements in the scene, and even the process of breaking a high-level task into a set of more concrete subtasks, a process called hierarchical task analysis, is environment-dependent. In this work, we propose ASHiTA, the first framework that generates a task hierarchy grounded to a 3D scene graph by breaking down high-level tasks into grounded subtasks. ASHiTA alternates LLM-assisted hierarchical task analysis, to generate the task breakdown, with task-driven 3D scene graph construction to generate a suitable representation of the environment. Our experiments show that ASHiTA performs significantly better than LLM baselines in breaking down high-level tasks into environment-dependent subtasks and is additionally able to achieve grounding performance comparable to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。