检测机器人示范中语言与动作不匹配的问题,提升训练数据质量。
Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations

- 将多模态信息视为专家,通过局部邻近和全局相似性估计任务标签分布。
- 在注入错误指令的LIBERO数据集上,检测准确率优于现有方法。
- 适用于需要语言澄清任务的机器人学习场景,可提升真实机器人性能。
用于训练视觉-语言-动作策略的机器人示范数据集可能包含一种隐蔽但有害的错误:行为正确但语言指令错误的轨迹。我们研究了此类指令-轨迹不匹配(ITM)的事后审计。与失败的推演不同,ITM通常看起来合理,却会污染策略所学的语言-行为映射。我们提出无需训练的多模态概率融合(MMPF)框架,将每种模态视为专家,基于局部邻域一致性与全局原型相似性估计任务标签分布,并通过预测熵加权在专家乘积中融合多模态信息。在注入指令错误的LIBERO基准和含噪声的真实机器人数据上,MMPF在ITM检测与标签修正方面表现最优。我们还证明,审计能提升多数下游策略学习效果,在需语言消歧的任务中尤为显著。真实机器人实验表明,该方法可提升策略性能,并揭示了过滤与重标注之间的权衡。
原文摘要 · Abstract (English)
Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with the wrong language instruction. We study post-hoc auditing of these Instruction-Trajectory Mismatches (ITMs). Unlike failed rollouts, ITMs often look plausible, and can corrupt the language-behavior mapping learned by the policy. We propose Multimodal Probabilistic Fusion (MMPF), a training-free auditing framework that treats each modality as an expert, estimates a task-label distribution from local neighborhood agreement and global prototype similarity, and then fuses modalities with predictive-entropy weighting in a product of experts. Across LIBERO benchmarks with injected instruction mismatches and noisy real-robot data, MMPF achieves the strongest overall ITM detection and label correction accuracy. We also show that auditing improves most downstream policy learning in settings where language is needed to disambiguate the task. We demonstrate in real robot experiments that our method can achieve improved policy performance and show the trade-off of filtering demonstrations compared to relabeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。