arXiv:2512.17958cs.ROcs.AI2025-12被引 1

用普通摄像头实现机器人实时理解人类意图,且在不同环境仍稳定有效。

Real-Time Human-Robot Interaction Intent Detection Using RGB-based Pose and Emotion Cues with Cross-Camera Model Generalization

  • 融合骨骼姿态与表情特征,仅靠单目摄像头和树莓派就能运行。
  • 在跨人、跨场景测试中,帧级和序列级准确率均达95%以上。
  • 适合部署在真实公共场所的低成本智能机器人,尤其看重鲁棒性。

公共空间中的服务机器人需要实时理解人类行为意图以实现自然交互。本文提出一种实用的多模态框架,通过单目RGB视频提取与相机无关的2D骨骼姿态和面部情绪特征,实现帧级精确的人机交互意图检测。相比依赖RGB-D传感器或GPU加速的先前方法,本方案可在资源受限的嵌入式硬件(树莓派5,仅CPU)上运行。为解决自然人机交互数据集中严重的类别不平衡问题,我们引入MINT-RVAE(多模态循环变分自编码器用于意图序列生成),合成时间一致的姿势-情绪-标签序列以实现数据重平衡。在跨被试和跨场景协议下的离线评估中表现出强泛化性能,帧级与序列级AUROC均达0.95。关键的是,我们在MIRA机器人头部进行跨摄像头实测,该设备采用不同的内置RGB传感器,在未见的真实非受控环境中运行。尽管存在领域偏移,部署系统在32次实时交互试验中仍达到91%准确率与100%召回率。离线与实际部署表现高度一致,证实了所提多模态方法在跨传感器与跨环境下的鲁棒性,凸显其在广泛多媒体社交机器人中的适用性。

原文摘要 · Abstract (English)

Service robots in public spaces require real-time understanding of human behavioral intentions for natural interaction. We present a practical multimodal framework for frame-accurate human-robot interaction intent detection that fuses camera-invariant 2D skeletal pose and facial emotion features extracted from monocular RGB video. Unlike prior methods requiring RGB-D sensors or GPU acceleration, our approach resource-constrained embedded hardware (Raspberry Pi 5, CPU-only). To address the severe class imbalance in natural human-robot interaction datasets, we introduce a novel approach to synthesize temporally coherent pose-emotion-label sequences for data re-balancing called MINT-RVAE (Multimodal Recurrent Variational Autoencoder for Intent Sequence Generation). Comprehensive offline evaluations under cross-subject and cross-scene protocols demonstrate strong generalization performance, achieving frame- and sequence-level AUROC of 0.95. Crucially, we validate real-world generalization through cross-camera evaluation on the MIRA robot head, which employs a different onboard RGB sensor and operates in uncontrolled environments not represented in the training data. Despite this domain shift, the deployed system achieves 91% accuracy and 100% recall across 32 live interaction trials. The close correspondence between offline and deployed performance confirms the cross-sensor and cross-environment robustness of the proposed multimodal approach, highlighting its suitability for ubiquitous multimedia-enabled social robots.

人机交互意图识别多模态融合嵌入式部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。