用生成视频提升机器人预判人类交接意图的准确率
RGB-D Video Generation for Improving Human-to-Robot Object Handover Prediction

- 用扩散模型生成带表情和姿态的RGB-D视频
- 生成视频让机器人零样本迁移时提前3秒识别交接意图
- 专为真实传感器噪声设计,适合实际机器人部署
人机交接是人机协作的核心能力,但受限于缺乏大规模、以人为核心的多模态数据集及显著的仿真到现实差距。为此,我们构建了Hand2Bot RGB-D视频数据集,包含人体姿态与面部表情等丰富上下文信息,并模拟真实传感器噪声。进一步提出PassGen生成框架,结合稳定视频扩散模型与意图感知时序人脸编码器,生成符合手物一致性的逼真交接序列。为弥合仿真与现实差距,引入基于形态的深度图编辑策略,复现真实深度图中的噪声模式。实验表明,该框架在消融测试与物理机器人部署中均实现高意图识别准确率和低误触发率。结果验证:基于PassGen训练可实现稳健的零样本迁移与更早意图预判,有效支持共享工作空间中的社交化机器人行为。
原文摘要 · Abstract (English)
Human-to-robot (H2R) object handover is a fundamental capability for human-robot collaboration, yet progress is hindered by the scarcity of large-scale, human-centric datasets and the significant sim-to-real gap. To address these challenges, we introduce Hand2Bot, an RGB-D video dataset that provides rich contextual information such as body posture and facial expressions, specifically collected for handover scenarios with real-world noise patterns. We further propose PassGen, a generative pipeline that leverages stable video diffusion and an Intention-Aware Temporal Face Encoder to synthesize realistic handover sequences while ensuring hand-object consistency. To bridge the sim-to-real gap, we implement a morphology-based depth editing strategy that replicates realistic sensor noise found in physical depth maps. Experimental evaluations demonstrate that our framework achieves high intention identification accuracy and low false trigger rates in both ablation studies and real-world deployment on a physical robot platform. Our results confirm that training on PassGen allows for robust zero-shot transfer and earlier intention anticipation compared to traditional hand-centric baselines, effectively enabling socially aware robotic behavior in shared workspaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。