arXiv:2509.22573cs.ROcs.CV2025-09被引 1

仅用摄像头图像预测人机交互意图,精度更高且响应更快。

MINT-RVAE: Multi-Cues Intention Prediction of Human-Robot Interaction using Human Pose and Emotion Information from RGB-only Camera Data

  • 基于RGB图像构建多线索意图预测模型,支持帧级精准判断。
  • 在真实数据集上达0.95的AUROC,优于已有方法的0.90-0.912。
  • 解决数据不平衡问题,开源新标注数据集供研究使用。

高效检测人类与普适机器人交互意图对实现有效人机协作至关重要。过去十年中,深度学习在该领域广泛应用,多数方法依赖多模态输入(如RGB-D)将时序传感数据划分为交互或非交互窗口。本文提出一种新颖的仅含RGB图像的端到端流水线,实现帧级精度的交互意图预测,提升机器人响应速度与服务质量。真实场景中人机交互数据存在类别不平衡问题,影响模型训练与泛化能力。为此,我们设计MINT-RVAE合成序列生成方法,结合新型损失函数与训练策略,显著增强模型在外部样本上的泛化性能。实验表明,本方法在关键指标上达到0.95的AUROC,优于先前工作(0.90–0.912),且仅需单目视觉输入。为推动后续研究,我们公开发布带帧级标注的新数据集。

原文摘要 · Abstract (English)

Efficiently detecting human intent to interact with ubiquitous robots is crucial for effective human-robot interaction (HRI) and collaboration. Over the past decade, deep learning has gained traction in this field, with most existing approaches relying on multimodal inputs, such as RGB combined with depth (RGB-D), to classify time-sequence windows of sensory data as interactive or non-interactive. In contrast, we propose a novel RGB-only pipeline for predicting human interaction intent with frame-level precision, enabling faster robot responses and improved service quality. A key challenge in intent prediction is the class imbalance inherent in real-world HRI datasets, which can hinder the model's training and generalization. To address this, we introduce MINT-RVAE, a synthetic sequence generation method, along with new loss functions and training strategies that enhance generalization on out-of-sample data. Our approach achieves state-of-the-art performance (AUROC: 0.95) outperforming prior works (AUROC: 0.90-0.912), while requiring only RGB input and supporting precise frame onset prediction. Finally, to support future research, we openly release our new dataset with frame-level labeling of human interaction intent.

人机交互意图预测视觉感知数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。