用智能眼镜麦克风阵列实现定向语音识别与干扰抑制
Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
- 通过序列化方向输出训练增强模型对声源方向的理解
- 在多说话人场景下实现高精度语音识别与声源定位
- 适合智能穿戴设备语音交互场景
近期研究证明,将音频编码输入大语言模型可实现有效的语音识别。然而,语音大模型处理带空间线索的多通道音频的能力仍缺乏深入研究。本文提出 directional-SpeechLlama,利用智能眼镜麦克风阵列实现定向语音识别、声源定位及旁听者语音干扰抑制。为提升模型对方向性的理解,提出两种关键技术:序列化方向输出训练(S-DOT)和对比方向数据增强(CDDA)。实验结果表明,该模型能有效捕捉文本提示与空间音频之间的关系,在语音识别与声源定位任务中均表现出色。
原文摘要 · Abstract (English)
Recent studies have demonstrated that prompting large language models (LLM) with audio encodings enables effective speech recognition capabilities. However, the ability of Speech LLMs to comprehend and process multi-channel audio with spatial cues remains a relatively uninvestigated area of research. In this work, we present directional-SpeechLlama, a novel approach that leverages the microphone array of smart glasses to achieve directional speech recognition, source localization, and bystander cross-talk suppression. To enhance the model's ability to understand directivity, we propose two key techniques: serialized directional output training (S-DOT) and contrastive direction data augmentation (CDDA). Experimental results show that our proposed directional-SpeechLlama effectively captures the relationship between textual cues and spatial audio, yielding strong performance in both speech recognition and source localization tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。