arXiv:2512.01022cs.RO2025-12中稿 · CVPR被引 7

提出新框架让机器人高效完成重复性操作任务

CycleManip: Enabling Cyclic Task Manipulation via Effective Historical Perception and Understanding

  • 通过成本感知采样与多任务学习增强历史记忆理解
  • 在仿真与真实场景中均实现高成功率,适配多种机械臂
  • 构建首个循环操作基准与自动评估工具,推动领域发展

本文探索机器人操控中一种重要但未被充分研究的任务:基于周期的操控,即机器人需在预期时间内完成循环或重复动作。这类任务在日常生活中至关重要,如摇晃瓶子或敲击钉子。然而,现有工作较少关注此类任务,导致两大挑战:1)模仿学习方法因历史信息利用不充分,常无法在预期时间内完成任务;2)缺乏足够数据和自动评估工具的基准,阻碍有效解决方案的发展。为此,我们提出CycleManip框架,以端到端模仿方式实现周期性任务操控,无需额外模型、分层结构或显著计算开销。核心思路是通过成本感知采样策略提升历史感知效率,并通过多任务学习增强历史理解。其次,我们构建了一个基于周期任务的基准,包含多样化的周期任务及自动评估方法。大量仿真与真实世界实验表明,所提方法在周期任务操控中取得高成功率,且在通用操控任务中表现出强适应性,可无缝集成于视觉-语言-动作(VLA)等模仿策略中。此外,该方法适用于多种机器人平台,包括双臂夹持器、灵巧手与人形机器人。

原文摘要 · Abstract (English)

In this paper, we explore an important yet underexplored task in robot manipulation: cycle-based manipulation, where robots need to perform cyclic or repetitive actions with an expected terminal time. These tasks are crucial in daily life, such as shaking a bottle or knocking a nail. However, few prior works have explored this task, leading to two main challenges: 1) the imitation methods often fail to complete these tasks within the expected terminal time due to the ineffective utilization of history; 2) the absence of a benchmark with sufficient data and automatic evaluation tools hinders development of effective solutions in this area. To address these challenges, we first propose the CycleManip framework to achieve cycle-based task manipulation in an end-to-end imitation manner without requiring any extra models, hierarchical structure or significant computational overhead. The core insight is to enhance effective history perception by a cost-aware sampling strategy and to improve historical understanding by multi-task learning. Second, we introduce a cycle-based task manipulation benchmark, which provides diverse cycle-based tasks, and an automatic evaluation method. Extensive experiments conducted in both simulation and real-world settings demonstrate that our method achieves high success rates in cycle-based task manipulation. The results further show strong adaptability performance in general manipulation, and the plug-and-play ability on imitation policies such as Vision-Language-Action (VLA) models. Moreover, the results show that our approach can be applied across diverse robotic platforms, including bi-arm grippers, dexterous hands, and humanoid robots.

机器人操控周期任务模仿学习多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。