将文本控制的抓取检测模型高效迁移到语音输入,提升人形机器人自然交互能力。
Speech2Grasp: Data-Efficient Transfer of Text-Conditioned Grasp Detection to Speech in Humanoid Robots

- 用轻量级MLP投影器将文本模型适配语音输入。
- 在真实机器人上性能超越语音转文字+文本模型的串联流程。
- 降低推理延迟,适合资源受限的实时机器人应用。
人形机器人需具备多模态理解能力以实现与人类的自然交互。尽管视觉-语言模型广泛应用,但通常假设输入为文本而非更自然的语音。本文研究是否可高效地将成熟的文本条件模型迁移到语音输入。以ALBEF为例,诊断分析表明,轻量级MLP投影器能有效适配其至语音,同时保持语义区分度和鲁棒性。基于此,提出Speech2Grasp框架,实现文本条件抓取检测到语音的高效迁移。真实人形机器人实验显示,该方法优于级联的语音识别+文本模型流程,且推理延迟更低。结果表明,这是一种可推广的将现有文本模型扩展至语音的实用范式。
原文摘要 · Abstract (English)
Humanoid robots increasingly require multi-modal understanding for natural interaction with humans. Despite the prominence of vision-language models, they generally assume textual rather than the more natural speech inputs. In this paper, we investigate whether a well-established text-conditioned model can be transferred to speech in a data-efficient manner. Using ALBEF as a case study, we conduct diagnostic analyses showing that a lightweight MLP-based projector effectively adapts it to speech, while preserving semantic discrimination and robustness. Motivated by these findings, we introduce Speech2Grasp, a framework for data-efficient transfer of text-conditioned grasp detection to speech. Real-world humanoid robot experiments show that Speech2Grasp outperforms cascaded ASR-based pipeline, while reducing inference latency. Our findings suggest a practical paradigm for extending established text-conditioned systems to speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。