arXiv:2606.08341cs.RO2026-06

用人类示范数据训练机器人,实现更可靠的动作意图预测与纠错。

Uncertainty-Aware Intention Prediction for Human-to-Robot Assembly Teleoperation

论文配图:Uncertainty-Aware Intention Prediction for Human-to-Robot Assembly Teleoperation
图 1 · 摘自论文原文
  • 用人类手势数据预训练,再少量机器人数据微调,提升动作理解能力。
  • 引入置信度评估机制,使系统在不确定时提前预警并纠正错误。
  • 适合需要高可靠性的人机协作装配任务,尤其适用于数据少的场景。

在人机协同装配的辅助遥操作中,准确预测用户意图对实现及时可靠的机器人协助至关重要。这类系统需持续理解用户行为,以识别动作、预测意图并实时检测错误。然而,机器人操作演示成本高且受硬件限制,而人类演示易于收集且富含时间结构。为此,我们提出一种不确定性感知的人机意图预测框架:(1) 分层迁移学习,先在人类手部动作数据上预训练MS-TCN++,再在少量机器人遥操作数据上微调,捕捉低级动作与高级任务意图;(2) 采用置信区间预测模块,提供帧级预测集合,并具备统计覆盖保证,实现可靠的不确定性量化与早期意图估计;(3) 基于视觉语言模型(VLM)的片段修正,利用视觉和时间上下文选择性审查低置信或时间不确定的片段。该框架支持动作识别、时间分割、意图预测与错误检测。在含22类动作的机器人装配演示上,仅用16次机器人示范,人到机器人的微调将测试集编辑分数从70.50提升至80.70。带安全保护的VLM修正进一步将帧准确率从45.21%提升至46.42%,同时提高F1@25与F1@50,且保持编辑分数不变。结果表明,人类示范可作为可扩展的预训练数据,支持鲁棒、不确定性感知的机器人动作分割。

原文摘要 · Abstract (English)

In assisted teleoperation for human-robot collaboration, accurate intention prediction is critical for enabling timely and reliable robotic assistance during long-horizon manipulation and assembly tasks. These systems require continuous understanding of user behavior to recognize actions, anticipate intentions, and detect mistakes in real time. However, robot teleoperation demonstrations are costly and hardware-limited, whereas human demonstrations are easier to collect and provide rich temporal structure. To address this challenge, we propose an uncertainty-aware human-to-robot intention prediction framework that combines: (1) hierarchical transfer learning, where MS-TCN++ is pretrained on human hand demonstrations and fine-tuned on limited robot teleoperation data to capture low-level actions and high-level task intentions; (2) a conformal prediction module that provides frame-level prediction sets with statistical coverage guarantees for reliable uncertainty quantification and early intention estimation; and (3) VLM-guided segment correction, which selectively reviews low-confidence or temporally uncertain segments using visual and temporal context. The framework supports action recognition, temporal segmentation, intention anticipation, and mistake detection for assisted teleoperation. Experiments on robot assembly demonstrations with 22 action classes show that human-to-robot fine-tuning improves the robot test-set Edit score from 70.50 to 80.70 using only 16 robot demonstrations. Edit-safe VLM correction further improves frame accuracy from 45.21% to 46.42% and increases F1@25 and F1@50 while preserving the Edit score. These results show that human demonstrations provide scalable pretraining data for robust, uncertainty-aware robot action segmentation. Code and data: project website.

意图预测人机协作不确定性建模遥操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。