arXiv:2502.02772cs.ROcs.AI2025-02被引 1

让机器人理解语言和力觉信号的联合表示,实现自然人机协作。

Cross-modality Force and Language Embeddings for Natural Human-Robot Communication

  • 构建语言与力觉信号的统一嵌入空间,实现跨模态对齐。
  • 在该空间中,语言与力觉可互补、融合或相互替代。
  • 适合研究人机交互、具身智能与多模态学习的研究者。

本文提出一种跨模态嵌入方法,将力觉特征与语言信息统一映射到同一潜在空间,实现言语与触觉通信的协同协调。当两人共同搬运重物时,会通过语言交流动作意图和施加的力,这种自然融合的语言与物理信号能有效促进协作。类似地,通过整合语言与力觉模态,人类-机器人交互也能达到同等水平的协调性。实验表明,尽管语言与力觉看似完全不同,但可在统一潜在空间中进行嵌入,并量化两者之间的语义相关性。在此空间中,力觉信号与语言可实现:(a) 相互补充,(b) 融合各自效果,(c) 以可交换方式互替。论文首先阐明跨模态嵌入的需求,介绍基础架构与关键技术模块,讨论数据采集方法与实现挑战,并展示实验结果与分析。

原文摘要 · Abstract (English)

A method for cross-modality embedding of force profile and words is presented for synergistic coordination of verbal and haptic communication. When two people carry a large, heavy object together, they coordinate through verbal communication about the intended movements and physical forces applied to the object. This natural integration of verbal and physical cues enables effective coordination. Similarly, human-robot interaction could achieve this level of coordination by integrating verbal and haptic communication modalities. This paper presents a framework for embedding words and force profiles in a unified manner, so that the two communication modalities can be integrated and coordinated in a way that is effective and synergistic. Here, it will be shown that, although language and physical force profiles are deemed completely different, the two can be embedded in a unified latent space and proximity between the two can be quantified. In this latent space, a force profile and words can a) supplement each other, b) integrate the individual effects, and c) substitute in an exchangeable manner. First, the need for cross-modality embedding is addressed, and the basic architecture and key building block technologies are presented. Methods for data collection and implementation challenges will be addressed, followed by experimental results and discussions.

人机交互跨模态力觉感知语言嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。