用视频光流提升机器人动作精度,解决语言指令操控中的低级动作不准问题。
ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow
- 将机器人动作建模为视频中的动作光流,自监督学习增强估计
- 通过迭代检索与去噪,实现动作光流的逐级精炼
- 动态记忆池融合历史光流,适合长时序视觉任务的高精度操控
语言指令驱动的机器人操作因从数据中学习的潜力而受到广泛关注。尽管高层感知与规划在通用大模型进展下持续改善,但低层动作估计的精度不足已成为操作性能的关键瓶颈。为此,本文提出一种新型机器人操作框架ActionSink,旨在推动基于学习的机器人操作中的精确动作估计。如其名称所示,ActionSink以自监督方式将机器人动作重构为视频中由动作引起的光流(称作“动作光流”),并将其检索与整合以增强动作估计。具体而言,ActionSink包含两个核心模块:第一是粗到细的动作光流匹配器,通过迭代检索与去噪过程持续提升动作光流的准确性;第二是动态动作光流集成器,采用工作记忆池动态高效管理应被用于当前动作估计的历史动作光流。该模块设计多层融合机制,集成当前直接估计值及来自当前与工作记忆的光流信息,通过一系列估计-融合过程实现高精度动作估计。所提ActionSink在LIBERO基准上较之前最先进方法提升7.9%的成功率,在具有挑战性的长时序视觉任务LIBERO-Long上获得近8%的准确率增益。
原文摘要 · Abstract (English)
Language-instructed robot manipulation has garnered significant interest due to the potential of learning from collected data. While the challenges in high-level perception and planning are continually addressed along the progress of general large pre-trained models, the low precision of low-level action estimation has emerged as the key limiting factor in manipulation performance. To this end, this paper introduces a novel robot manipulation framework, i.e., ActionSink, to pave the way toward precise action estimations in the field of learning-based robot manipulation. As the name suggests, ActionSink reformulates the actions of robots as action-caused optical flows from videos, called "action flow", in a self-supervised manner, which are then used to be retrieved and integrated to enhance the action estimation. Specifically, ActionSink incorporates two primary modules. The first module is a coarse-to-fine action flow matcher, which continuously refines the accuracy of action flow via iterative retrieval and denoising process. The second module is a dynamic action flow integrator, which employs a working memory pool that dynamically and efficiently manages the historical action flows that should be used to integrate to enhance the current action estimation. In this module, a multi-layer fusion module is proposed to integrate direct estimation and action flows from both the current and the working memory, achieving highly accurate action estimation through a series of estimation-integration processes. Our ActionSink framework outperformed prior SOTA on the LIBERO benchmark by a 7.9\% success rate, and obtained nearly an 8\% accuracy gain on the challenging long-horizon visual task LIBERO-Long.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。