arXiv:2604.23000cs.RO2026-04被引 1

用运动平滑性筛选演示数据,提升模仿学习效果。

Learning from the Best: Smoothness-Driven Metrics for Data Quality in Imitation Learning

论文配图:Learning from the Best: Smoothness-Driven Metrics for Data Quality in Imitation Learning
图 1 · 摘自论文原文
  • 基于轨迹平滑性设计评分机制,无需训练或标注。
  • 仅用1/6数据达16%成功率提升,半数数据提20%表现。
  • 适用于数据清洗、检索重排与领域加权,适合真实场景。

在行为克隆中,策略性能受限于示范数据质量。现实数据因操作者技能差异、遥操作噪声和流程不一致导致轨迹质量参差不齐,但传统方法对所有示范同等对待。现有清理方法需依赖政策训练循环或人工标注,难以扩展。本文提出RINSE(Ranking and INdexing Smooth Examples),一种轻量级框架,仅通过轨迹数据计算平滑性得分,且不依赖策略结构。结合电机控制理论,采用两个互补指标:频域正则性度量Spectral Arc Length(SAL)和接触感知空间偏差度量Trajectory-Envelope Distance(TED)。实验证明,平滑性过滤可降低保留数据的条件动作方差,且该效应可通过动作分块放大。在RoboMimic基准上,使用六分之一数据时,SAL过滤使成功率提升16%;在真实操作任务中,使用一半数据时,TED过滤实现20%性能提升。在LIBERO-10上的STRAP框架中,RINSE重排序使平均成功率提升5.6%。作为Re-Mix领域重加权中的软权重,其得分与学习到的分配高度相关(斯皮尔曼ρ≥0.89)。结果表明,平滑性是过滤、检索与重加权场景下的有效质量信号,尤其适用于噪声或异构数据环境。

原文摘要 · Abstract (English)

In behavioral cloning (BC), policy performance is fundamentally limited by demonstration data quality. Real-world datasets contain trajectories of varying quality due to operator skill differences, teleoperation artifacts, and procedural inconsistencies, yet standard BC treats all demonstrations equally. Existing curation methods require costly policy training in the loop or manual annotation, limiting scalability. We propose RINSE (Ranking and INdexing Smooth Examples), a lightweight framework for scoring demonstrations based on trajectory smoothness that is policy-architecture-agnostic and operates on trajectory data alone, with TED additionally using a phase-boundary/contact signal. Grounded in motor control theory, which establishes smoothness as a hallmark of skilled movement, RINSE uses two complementary metrics: Spectral Arc Length (SAL), a spectral measure of frequency-domain regularity, and Trajectory-Envelope Distance (TED), a spatial measure of contact-aware geometric deviation. We show that smoothness filtering can reduce the conditional action variance of the retained data distribution, with downstream effects that can be amplified by action chunking and compounding error. On RoboMimic benchmarks, SAL filtering achieves 16% higher success using one-sixth of the data. On real-world manipulation, TED filtering achieves 20% improvement with half the data. As a retrieval-stage filter within STRAP on LIBERO-10, RINSE re-ranking improves mean success by 5.6%. As soft weights in Re-Mix domain reweighting, RINSE scores produce domain allocations highly correlated with the learned Re-Mix allocations (Spearman $ρ\geq 0.89$). These results support smoothness as a useful quality signal across filtering, retrieval, and reweighting settings, especially in noisy or heterogeneous data regimes.

模仿学习数据质量平滑性机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。