让智能体同时做饭清洁,用调度优化提升效率
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution
- 引入基于运筹学的3D任务调度新范式,支持并行执行子任务
- 构建6万条真实场景任务数据集,最小化总完成时间
- 提出GRANT模型,融合语言理解与空间定位实现高效调度
任务调度对具身智能至关重要,使智能体能在3D物理世界中理解自然语言指令并高效执行动作。然而,现有数据集常忽略运筹学(OR)知识与3D空间定位。本文提出基于运筹学知识的3D grounded任务调度(ORS3D),要求智能体协同语言理解、3D定位与效率优化,通过并行执行子任务(如边加热边清洁水槽)来最小化总完成时间。为推动研究,我们构建了包含6万条复合任务的大型数据集ORS3D-60K,覆盖4000个真实场景。同时提出GRANT——一种具身多模态大模型,采用简单有效的调度标记机制生成高效任务序列与空间动作。在ORS3D-60K上的大量实验验证了GRANT在语言理解、3D定位和调度效率方面的有效性。
原文摘要 · Abstract (English)
Task scheduling is critical for embodied AI, enabling agents to follow natural language instructions and execute actions efficiently in 3D physical worlds. However, existing datasets often simplify task planning by ignoring operations research (OR) knowledge and 3D spatial grounding. In this work, we propose Operations Research knowledge-based 3D Grounded Task Scheduling (ORS3D), a new task that requires the synergy of language understanding, 3D grounding, and efficiency optimization. Unlike prior settings, ORS3D demands that agents minimize total completion time by leveraging parallelizable subtasks, e.g., cleaning the sink while the microwave operates. To facilitate research on ORS3D, we construct ORS3D-60K, a large-scale dataset comprising 60K composite tasks across 4K real-world scenes. Furthermore, we propose GRANT, an embodied multi-modal large language model equipped with a simple yet effective scheduling token mechanism to generate efficient task schedules and grounded actions. Extensive experiments on ORS3D-60K validate the effectiveness of GRANT across language understanding, 3D grounding, and scheduling efficiency. The code is available at https://github.com/H-EmbodVis/GRANT
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。