将波斯语唇读技术融入机器人,提升嘈杂环境下的交互能力
Integrating Persian Lip Reading in Surena-V Humanoid Robot for Human-Robot Interaction
- 用关键点追踪和深度网络双方法实现波斯语唇读
- LSTM模型达89%准确率,已实现在机器人上实时运行
- 适合语音受限场景的护理与服务类机器人应用
唇读对社交环境中的人机交互至关重要,可增强机器人在嘈杂环境中的沟通能力,尤其适用于照护和客服场景。本研究构建了首个波斯语唇读数据集,并将唇读技术集成至Surena-V人形机器人,以提升其语音识别能力。采用两种互补方法:一是基于面部关键点追踪的间接法,聚焦嘴唇区域运动推断;二是直接使用卷积神经网络(CNN)与长短期记忆(LSTM)网络处理原始视频流进行动作与语音识别。其中,LSTM模型表现最佳,达到89%的准确率,并成功部署于Surena-V机器人中,实现了实时人机交互。研究表明,该方法在语音受限环境下具有显著有效性。
原文摘要 · Abstract (English)
Lip reading is vital for robots in social settings, improving their ability to understand human communication. This skill allows them to communicate more easily in crowded environments, especially in caregiving and customer service roles. Generating a Persian Lip-reading dataset, this study integrates Persian lip-reading technology into the Surena-V humanoid robot to improve its speech recognition capabilities. Two complementary methods are explored, an indirect method using facial landmark tracking and a direct method leveraging convolutional neural networks (CNNs) and long short-term memory (LSTM) networks. The indirect method focuses on tracking key facial landmarks, especially around the lips, to infer movements, while the direct method processes raw video data for action and speech recognition. The best-performing model, LSTM, achieved 89\% accuracy and has been successfully implemented into the Surena-V robot for real-time human-robot interaction. The study highlights the effectiveness of these methods, particularly in environments where verbal communication is limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。