arXiv:2608.04196cs.ROcs.CV2026-08

用相似性筛选3200万个人类视频,仅用149万样本提升机器人操作成功率。

SiMDex: Mining Similar Egocentric Videos for Cross-Embodiment Dexterous Manipulation

论文配图:SiMDex: Mining Similar Egocentric Videos for Cross-Embodiment Dexterous Manipulation
图 1 · 摘自论文原文
  • 将人类视频筛选视为推荐问题,三阶段管道精准提取相关数据。
  • 仅用149万样本(<5%)就使成功率从47.7%提升至61.1%。
  • 无需修改模型结构,适用于任意形态的机器人操作任务。

近年来,以第一人称视角的人类视频在机器人操作中快速增长,但尚不清楚哪些数据真正有助于精细操作。我们提出SiMDex,一种基于相似性的数据挖掘框架,将视觉-语言-动作模型(VLA)后训练中的人类数据选择问题转化为推荐问题。针对每个机器人示范,SiMDex采用三层召回-排序-重排序流程,从约3200万条第一人称人类样本中提取任务相关子集,在不依赖具体身体形态的动作空间中运行,无需修改VLA架构或训练流程。与使用等量随机采样人类数据的强基线相比,SiMDex仅用约149万样本(低于数据池的5%),即把整体成功率从47.7%提升至61.1%,证明了有选择的数据筛选优于盲目混合。

原文摘要 · Abstract (English)

Recent years have witnessed an explosive trend of scaling ego-centric human videos for robot manipulation, yet it remains unclear which data actually benefits dexterous manipulation. We present SiMDex, a similarity-based data mining framework that casts human data selection for VLA post-training in dexterous manipulation as a recommendation problem. For each robot demonstration, SiMDex employs a three-layer recall-ranking-re-ranking pipeline to extract task-relevant subsets from a pool of ~32M egocentric human samples, operating in a morphology-agnostic action space that requires no changes to VLA architecture or training. Against a strong baseline trained with an equal amount of randomly sampled human data, SiMDex uses only ~1.49M mined samples (<5% of the pool) yet improves the overall success rate from 47.7% to 61.1%, showing that selective curation outperforms indiscriminate data mixing.

数据挖掘机器人操作视觉语言模型第一人称视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。