arXiv:2512.07371cs.ROcs.AI2025-12被引 7

用语义感知方法压缩演示数据,让机器人模仿人类动作更快更准。

ESPADA: Execution Speedup via Semantics Aware Demonstration Data Downsampling for Imitation Learning

  • 通过视觉语言模型识别动作关键段,仅对非关键部分大幅降采样。
  • 在仿真与真实场景中实现约2倍加速,成功率几乎不变。
  • 无需额外数据或重训练,适合想提升机器人执行速度的研究者。

基于行为克隆的视觉运动策略虽能实现精准操作,但常继承人类示范的缓慢、谨慎节奏,限制实际部署。以往加速方法多依赖统计或启发式线索,忽视任务语义,在不同操作场景中易失效。本文提出ESPADA,一种语义与空间感知框架,利用视觉语言模型-大语言模型(VLM-LLM)流水线结合3D抓取器-物体关系,对演示数据进行分段,仅在非关键段进行激进降采样,同时保留关键阶段精度,无需额外数据、架构修改或重训练。为将单条标注轨迹扩展至全数据集,ESPADA基于仅动力学特征的动态时间规整(DTW)传播分段标签。在ACT与DP基准上,无论仿真还是真实世界实验,ESPADA均实现约2倍速度提升,同时保持成功率,显著缩小人类示范与高效机器人控制之间的差距。

原文摘要 · Abstract (English)

Behavior-cloning based visuomotor policies enable precise manipulation but often inherit the slow, cautious tempo of human demonstrations, limiting practical deployment. However, prior studies on acceleration methods mainly rely on statistical or heuristic cues that ignore task semantics and can fail across diverse manipulation settings. We present ESPADA, a semantic and spatially aware framework that segments demonstrations using a VLM-LLM pipeline with 3D gripper-object relations, enabling aggressive downsampling only in non-critical segments while preserving precision-critical phases, without requiring extra data or architectural modifications, or any form of retraining. To scale from a single annotated episode to the full dataset, ESPADA propagates segment labels via Dynamic Time Warping (DTW) on dynamics-only features. Across both simulation and real-world experiments with ACT and DP baselines, ESPADA achieves approximately a 2x speed-up while maintaining success rates, narrowing the gap between human demonstrations and efficient robot control.

模仿学习速度加速语义感知机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。