arXiv:2512.16302cs.RO2025-12AAAI被引 3

仅用一次示范教机器人完成复杂长时序抓取任务

ManiLong-Shot: Interaction-Aware One-Shot Imitation Learning for Long-Horizon Manipulation

  • 将长任务拆解为交互事件驱动的原子动作
  • 仅10个短任务示范即可泛化到20个新长任务
  • 适合快速部署复杂抓取操作的机器人系统

单次示范模仿学习(OSIL)为无需大规模数据收集即可教会机器人新技能提供了可能。然而,现有方法主要局限于短时序任务,难以应对复杂的长时序操作。为此,我们提出ManiLong-Shot框架,实现长时序抓取任务的有效单次示范学习。该框架以物理交互事件为基础结构长时序任务,将问题重构为序列化交互感知的原子动作,而非直接模仿连续轨迹。这一分解可由视觉语言模型(VLM)的高层推理或基于机器人状态变化的规则启发式方法驱动。针对每个原子动作,ManiLong-Shot预测交互关键区域,建立示范与当前观测间的对应关系,并计算目标末端执行器位姿,从而实现有效任务执行。大量仿真实验表明,仅在10个短时序任务上训练后,该方法即可通过单次示范泛化至三个难度等级下的20个未见长时序任务,相对最先进方法提升22.8%。此外,真实机器人实验验证了ManiLong-Shot在单次示范下稳健执行三个长时序操作的能力,证实其实际应用可行性。

原文摘要 · Abstract (English)

One-shot imitation learning (OSIL) offers a promising way to teach robots new skills without large-scale data collection. However, current OSIL methods are primarily limited to short-horizon tasks, thus limiting their applicability to complex, long-horizon manipulations. To address this limitation, we propose ManiLong-Shot, a novel framework that enables effective OSIL for long-horizon prehensile manipulation tasks. ManiLong-Shot structures long-horizon tasks around physical interaction events, reframing the problem as sequencing interaction-aware primitives instead of directly imitating continuous trajectories. This primitive decomposition can be driven by high-level reasoning from a vision-language model (VLM) or by rule-based heuristics derived from robot state changes. For each primitive, ManiLong-Shot predicts invariant regions critical to the interaction, establishes correspondences between the demonstration and the current observation, and computes the target end-effector pose, enabling effective task execution. Extensive simulation experiments show that ManiLong-Shot, trained on only 10 short-horizon tasks, generalizes to 20 unseen long-horizon tasks across three difficulty levels via one-shot imitation, achieving a 22.8% relative improvement over the SOTA. Additionally, real-robot experiments validate ManiLong-Shot's ability to robustly execute three long-horizon manipulation tasks via OSIL, confirming its practical applicability.

机器人学习单次示范长时序操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。