arXiv:2606.24403cs.ROcs.LG2026-06

用可解释的模块化框架,让机器人学会更懂物体动作的模仿技能。

RE4: Transformation-aware Imitation of Object Interactions Using Manipulation Modes

论文配图:RE4: Transformation-aware Imitation of Object Interactions Using Manipulation Modes
图 1 · 摘自论文原文
  • 基于姿态估计和操作模式检索,分四步重构示范轨迹。
  • 在图像和状态输入下,性能优于现有方法且鲁棒性更强。
  • 适合需要透明决策过程的机器人操作研究者使用。

物体交互任务是模仿学习的重要方向。尽管以扩散和流模型为主导的端到端方法性能显著提升,但牺牲了可解释性。本文重新审视了几种现代物体交互模仿学习基准,提出一个融合操作原理的框架,兼顾性能与可解释性。针对图像观测,利用演示数据中的自监督信号,实现轻量级无模型目标物体姿态估计。该信息用于:1)基于操作模式的示范检索;2)模式感知的变换;3)保持模式约束的重规划;4)滚动执行变换后的示范。该框架在Push-T和Robomimic的状态与图像基准上评估,对抗性基准显示其在稀疏数据区域的鲁棒性,低数据实验进一步验证其优势。结果表明,简单可解释的组件组合可有效学习操作技能。

原文摘要 · Abstract (English)

Object interaction tasks have been a focus of advances in imitation learning. End-to-end methods, dominated by diffusion and flow-based variants have shown leaps in performance while sacrificing interpretability. Object-centric and pose-informed variants have had a role in learning from demonstration in manipulation tasks. In this paper, we revisit a few modern imitation learning benchmarks for object interactions, with the aim of composing a framework that repurposes principled theories of manipulation, preserving both performance and interpretability. For image observations, lightweight training is proposed for model-free pose estimation of the target object, using self-supervision over the demonstration data available for imitation learning. This information is then used to inform a manipulation mode-aware retrieval of a demonstration, a mode-aware transformation, a replan step that connects to the retrieval point while preserving mode constraints, and finally rolling out the transformed demonstration. These compose four key steps of the proposed RE4 framework, evaluated over state-based and image-based benchmarks in Push-T and Robomimic. An adversarial benchmark that evaluates sparse data regions of image-based Push-T showcases the robustness, further bolstered by indications from low-data regime experiments. The current work shows promise in using simple interpretable building blocks to learn manipulation skills.

模仿学习机器人操作可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。