用语言统一控制机器人在动态环境中的动作与空间理解
ACTLLM: Action Consistency Tuned Large Language Model
- 用语言构建结构化场景描述,统一任务执行与空间推理
- 引入动作一致性约束,提升视觉表征的动作相关性
- 将任务建模为多轮视觉对话,增强长期任务的上下文关联
本文提出ACTLLM(Action Consistency Tuned Large Language Model),一种用于动态环境中机器人操作的新方法。传统视觉系统常难以同时具备优秀的任务执行与空间推理能力,限制了其在动态环境中的适应性。ACTLLM通过语言生成结构化场景描述,以灵活语言指令实现空间理解与任务性能的统一接口。此外,引入新的动作一致性约束,使视觉感知与对应动作对齐,从而增强可操作的视觉表征学习。同时,将操作任务的马尔可夫决策过程重构为多轮视觉对话框架,利用任务历史实现长期任务执行的上下文增强。评估表明,ACTLLM在多样场景中表现优异,验证了其在复杂视觉机器人操作任务中的有效性。
原文摘要 · Abstract (English)
This paper introduces ACTLLM (Action Consistency Tuned Large Language Model), a novel approach for robot manipulation in dynamic environments. Traditional vision-based systems often struggle to learn visual representations that excel in both task execution and spatial reasoning, thereby limiting their adaptability in dynamic environments. ACTLLM addresses these challenges by harnessing language to craft structured scene descriptors, providing a uniform interface for both spatial understanding and task performance through flexible language instructions. Moreover, we introduce a novel action consistency constraint that aligns visual perception with corresponding actions, thereby enhancing the learning of actionable visual representations. Additionally, we have reformulated the Markov decision process for manipulation tasks into a multi-turn visual dialogue framework. This approach enables the modeling of long-term task execution with enhanced contextual relevance derived from the history of task execution. During our evaluation, ACTLLM excels in diverse scenarios, proving its effectiveness on challenging vision-based robot manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。