让大模型学会听清智能眼镜中的多人对话方向
Equipping LLM with Directional Multi-Talker Speech Understanding Capabilities
- 用多麦克风阵列配合语音分离或端到端训练,实现方向感知
- 在语音识别和翻译任务中表现优异,支持实时流式处理
- 适合智能穿戴设备中的多说话人场景,如智能眼镜
近期研究显示,将音频编码输入大语言模型可实现有效的语音理解。然而,多数语音大模型仅在单通道、单说话人数据上训练,难以直接应用于多说话人与多通道语音理解任务。本文针对智能眼镜场景,系统研究如何为大模型赋予方向性多说话人语音理解能力。提出两种新方法:(1) 前端采用语音分离模块的级联系统;(2) 利用序列化输出训练的端到端系统。所有方法均基于智能眼镜内置的多麦克风阵列,实现方向性的实时解析与处理。实验结果表明,所提方法有效提升了大模型在语音识别与语音翻译任务中的性能,具备良好的实用性。
原文摘要 · Abstract (English)
Recent studies have demonstrated that prompting large language models (LLM) with audio encodings enables effective speech understanding capabilities. However, most speech LLMs are trained on single-channel, single-talker data, which makes it challenging to directly apply them to multi-talker and multi-channel speech understanding task. In this work, we present a comprehensive investigation on how to enable directional multi-talker speech understanding capabilities for LLMs, specifically in smart glasses usecase. We propose two novel approaches to integrate directivity into LLMs: (1) a cascaded system that leverages a source separation front-end module, and (2) an end-to-end system that utilizes serialized output training. All of the approaches utilize a multi-microphone array embedded in smart glasses to optimize directivity interpretation and processing in a streaming manner. Experimental results demonstrate the efficacy of our proposed methods in endowing LLMs with directional speech understanding capabilities, achieving strong performance in both speech recognition and speech translation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。