arXiv:2606.15587cs.RO2026-06被引 1

流畅演示反成教学负担,新方法从关键动作段提取强化信号。

Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments

论文配图:Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments
图 1 · 摘自论文原文
  • 从演示中提取关键运动片段,增强模型对精细操作的感知
  • 仅用流畅数据训练,性能逼近刻意示范(62.2% vs 64.4%)
  • 适合需要高精度操控的机器人学习任务

专家演示常被视为机器人模仿学习的黄金标准。但在插入、堆叠、对齐等细粒度操作中,我们发现一种反直觉现象:流畅的演示反而成为劣质教师。熟练操作员将对齐和恢复的关键时刻压缩在极短时间窗口内,导致策略被冗余的空域运动淹没,而精确决定成败处却缺乏监督。我们从数据与表征双层面解决此瓶颈。数据层面,放慢对齐阶段并重采样关键段落虽有帮助,但主要收益来自扩大恢复状态的覆盖范围,而非重新加权已有帧。然而,该方法未改变策略的逐帧视图:单张图像仍直接映射动作,修正所依赖的局部运动仍隐含。因此,我们转向表征层,提出STAIR(时空特征作为机器人学习接口),一种紧凑的动态特征,连接视觉语言模型与动作专家,将轨迹中已记录的短期运动提炼为密集、具运动感知的监督信号。仅在流畅数据上训练,STAIR即可实现50.0%到62.2%的整体性能,接近刻意演示的64.4%。这提示应以机器可学性为优化目标,而非仅追求人类效率。

原文摘要 · Abstract (English)

Expert demonstrations are widely assumed to be the gold standard for robot imitation learning. Yet for fine-grained manipulation such as insertion, stacking, and alignment, we uncover a counterintuitive failure mode: fluent demonstrations can be poor teachers. A skilled teleoperator compresses the decisive moments of alignment and recovery into a brief temporal window, leaving the policy flooded with redundant free-space motion and starved of supervision exactly where precision determines success. We address this bottleneck at two levels. At the data level, slowing down near alignment and resampling critical segments both help, yet the gain comes mainly from broadening the coverage of recovery states the policy must learn, not from reweighting frames it already has. Such data-side fixes, however, leave the policy's per-frame view untouched: a single image still maps directly to an action, and the local motion that governs correction stays implicit. We therefore turn to the representation level and introduce STAIR (\textbf{S}patio-\textbf{T}emporal feature \textbf{A}s an \textbf{I}nterface for \textbf{R}obot learning), a compact dynamic feature that bridges the vision-language model and the action expert, distilling the short-horizon motion already recorded in each trajectory into dense, motion-aware supervision. Trained on fluent data alone, STAIR recovers most of the deliberate-demonstration gain ($50.0$ to $62.2\%$ overall, approaching the $64.4\%$ of deliberate demonstrations). These results call for a more pedagogical view of robot data, optimized for machine learnability rather than human efficiency alone.

机器人学习模仿学习动作对齐动态特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。