让机器人根据任务自动选择空间表示方式,提升操作效率与稳定性。
T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models
- 按任务需求动态选择空间表示提取方案
- 实测在真实场景中提升理解力与运行效率
- 无需额外训练,适合实际部署的机器人系统
构建能在真实环境中完成多种任务的通用机器人操作体系仍具挑战性。视觉语言模型(VLMs)在机器人操作中展现出巨大潜力,主要得益于其从大规模数据集中获得的丰富世界知识。在此过程中,空间表示(如物体位置点或方向向量)充当VLMs与现实场景之间的桥梁,有效将VLM的推理能力落地到具体任务中。然而,现有基于VLM的机器人方法通常对各类任务采用固定的空表示提取方式,导致表征能力不足或提取耗时过长。本文提出T-Rex——一种任务自适应的空间表示提取框架,能根据具体任务需求动态选择最合适的表示提取策略。核心洞察是:任务复杂度决定空间表示的类型与粒度,更强的表征能力通常伴随更高的系统开销。在真实机器人环境中的大量实验表明,该方法在空间理解、效率和稳定性方面均有显著提升,且无需额外训练。
原文摘要 · Abstract (English)
Building a general robotic manipulation system capable of performing a wide variety of tasks in real-world settings is a challenging task. Vision-Language Models (VLMs) have demonstrated remarkable potential in robotic manipulation tasks, primarily due to the extensive world knowledge they gain from large-scale datasets. In this process, Spatial Representations (such as points representing object positions or vectors representing object orientations) act as a bridge between VLMs and real-world scene, effectively grounding the reasoning abilities of VLMs and applying them to specific task scenarios. However, existing VLM-based robotic approaches often adopt a fixed spatial representation extraction scheme for various tasks, resulting in insufficient representational capability or excessive extraction time. In this work, we introduce T-Rex, a Task-Adaptive Framework for Spatial Representation Extraction, which dynamically selects the most appropriate spatial representation extraction scheme for each entity based on specific task requirements. Our key insight is that task complexity determines the types and granularity of spatial representations, and Stronger representational capabilities are typically associated with Higher overall system operation costs. Through comprehensive experiments in real-world robotic environments, we show that our approach delivers significant advantages in spatial understanding, efficiency, and stability without additional training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。