arXiv:2605.27114cs.RO2026-05被引 1

用沉浸式VR高效收集机器人操作数据,只修正最不确定的失败片段。

VR-DAgger: Immersive VR for Dexterous Data Collection and Uncertainty-Guided On-Policy Correction

论文配图:VR-DAgger: Immersive VR for Dexterous Data Collection and Uncertainty-Guided On-Policy Correction
图 1 · 摘自论文原文
  • 通过VR实现直观操控,结合模拟环境自动回放
  • 仅对高不确定性失败片段进行标注,提升效率40%
  • 适合需要精细操作的数据收集场景

学习演示在机器人操作中有效,但任务特定数据的采集仍是主要瓶颈。分布偏移下,小误差累积导致性能下降,专家常花费大量时间处理重复性低价值修正。我们提出VR-DAgger,一个以沉浸式VR为核心的闭环框架,支持灵巧机械手的远程操作、演示数据采集与选择性策略修正。VR客户端提供同步场景可视化和直观手部控制,后端工作站运行仿真与学习,实现无需持续监督的自主回放。利用蒙特卡洛丢弃法评估扩散策略在Isaac Lab中的不确定性,筛选出有信息量的失败段落作为视频回放,在VR中由操作员选择性标注并修正行为,聚焦于高不确定性区域,无需全程监控或额外干预分类器。在三个灵巧操作任务(盘子抓取放置、抽屉打开、阀门旋转)上,使用10自由度XHand进行测试,涵盖标准与挑战性初始配置。主动标注在所有任务中均优于行为克隆,最高提升23个百分点。相比无引导的人机协同检查,VR-DAgger将每样本采集时间减少约40%,仅审查选定片段而非完整回放。

原文摘要 · Abstract (English)

Learning from demonstrations is effective for robotic manipulation, but collecting sufficient task-specific data remains a major bottleneck. Under distribution shift, small errors compound, performance degrades, and expert time is often spent on redundant, low-value corrections instead of the few critical failure cases. We present VR-DAgger, a human-in-the-loop framework centered on an immersive VR application for dexterous teleoperation, demonstration collection, and selective policy correction. The VR client provides intuitive hand control with synchronized scene visualization, while a backend workstation runs simulation and learning, enabling autonomous rollouts without continuous operator oversight. We use Monte Carlo (MC) dropout to score uncertainty during Isaac Lab rollouts of a diffusion policy and select informative failure segments for correction. These segments are replayed in VR as clips, where the operator selectively labels and corrects the policy's behavior, concentrating supervision where uncertainty is highest without full-rollout monitoring or a separate intervention classifier. We evaluate on three dexterous manipulation tasks (Pan pick-and-place, Drawer opening, Valve turning) with a 10-DoF XHand under standard and challenging initial configurations. Active labeling consistently improves over behavioral cloning across all tasks, with gains of up to 23 percentage points. Compared to unguided human-in-the-loop inspection, VR-DAgger reduces per-sample collection time by approximately 40% by focusing review on selected segments rather than full rollouts.

机器人操作虚拟现实强化学习数据收集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。