利用语音语调提升机器人对模糊指令的理解与执行能力
Enhancing Speech Instruction Understanding and Disambiguation in Robotics via Speech Prosody
- 直接分析语音语调而非仅转文本,捕捉关键语义线索
- 在机器人指令中识别目标意图准确率达95.79%,任务计划选择准确率71.96%
- 首个面向机器人语音歧义的研究数据集,适合人机交互方向研究者
让机器人准确理解并执行口头指令是实现高效人机协作的关键。传统方法依赖语音识别将语音转为文本,常忽略对意图消歧至关重要的语调信息。本文提出一种新方法,直接利用语音语调推断并解析指令意图,将预测的意图通过上下文学习融入大语言模型,以消歧并选择合适的任务规划。此外,我们构建了首个面向机器人领域的语音歧义数据集,推动该方向研究。实验表明,该方法在识别句内指代意图方面达到95.79%准确率,对模糊指令的任务计划选择准确率为71.96%,显著提升人机沟通效能。
原文摘要 · Abstract (English)
Enabling robots to accurately interpret and execute spoken language instructions is essential for effective human-robot collaboration. Traditional methods rely on speech recognition to transcribe speech into text, often discarding crucial prosodic cues needed for disambiguating intent. We propose a novel approach that directly leverages speech prosody to infer and resolve instruction intent. Predicted intents are integrated into large language models via in-context learning to disambiguate and select appropriate task plans. Additionally, we present the first ambiguous speech dataset for robotics, designed to advance research in speech disambiguation. Our method achieves 95.79% accuracy in detecting referent intents within an utterance and determines the intended task plan of ambiguous instructions with 71.96% accuracy, demonstrating its potential to significantly improve human-robot communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。