arXiv:2512.06558cs.RO2025-12中稿 · the ACM/IEEE Inter…被引 1

构建多视角人机交互数据集,提升机器人理解指令能力。

Embodied Referring Expression Comprehension in Human-Robot Interaction

  • 提出多模态引导残差模块,增强跨模态特征提取。
  • 在四个数据集上验证,模型性能显著提升。
  • 适合研究人机交互与具身语言理解的学者。

随着机器人进入人类工作空间,理解具身化的人类指令成为实现自然流畅人机交互的关键。然而,由于缺乏大规模、多样化的具身交互数据集,准确理解仍面临挑战。现有数据集存在视角偏差、单视角采集、非语言手势覆盖不足以及以室内环境为主等问题。为此,我们提出了Refer360数据集,该数据集在室内外多种场景中,从多视角收集了大规模的具身言语与非言语交互数据。同时,我们引入MuRes模块,一种多模态引导残差机制,通过作为信息瓶颈提取模态特异性信号,并将其强化至预训练表征中,生成互补特征用于下游任务。我们在四个HRI数据集(包括Refer360)上进行大量实验,发现当前多模态模型无法全面捕捉具身交互;但加入MuRes后,性能持续提升。这些结果确立了Refer360作为重要基准的价值,并展示了引导残差学习在提升机器人在人类环境中具身指代理解方面的潜力。

原文摘要 · Abstract (English)

As robots enter human workspaces, there is a crucial need for them to comprehend embodied human instructions, enabling intuitive and fluent human-robot interaction (HRI). However, accurate comprehension is challenging due to a lack of large-scale datasets that capture natural embodied interactions in diverse HRI settings. Existing datasets suffer from perspective bias, single-view collection, inadequate coverage of nonverbal gestures, and a predominant focus on indoor environments. To address these issues, we present the Refer360 dataset, a large-scale dataset of embodied verbal and nonverbal interactions collected across diverse viewpoints in both indoor and outdoor settings. Additionally, we introduce MuRes, a multimodal guided residual module designed to improve embodied referring expression comprehension. MuRes acts as an information bottleneck, extracting salient modality-specific signals and reinforcing them into pre-trained representations to form complementary features for downstream tasks. We conduct extensive experiments on four HRI datasets, including the Refer360 dataset, and demonstrate that current multimodal models fail to capture embodied interactions comprehensively; however, augmenting them with MuRes consistently improves performance. These findings establish Refer360 as a valuable benchmark and exhibit the potential of guided residual learning to advance embodied referring expression comprehension in robots operating within human environments.

人机交互具身理解多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。