多模态模型在噪声中提升语音识别,关键看模态类型和噪声程度。
MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
- 融合多种模态提供互补信息,提升语音识别准确率。
- 同步模态(如口型)在高噪声下更有效,非同步模态(如图像)在中等噪声下优势明显。
- 高质量视觉表示对识别性能有持续提升作用,适合噪声环境下的语音系统设计。
近年来,多模态大语言模型(MLLMs)为统一建模语音、文本、图像等模态提供了新可能。基于前期工作,本文研究了在何种条件下及何种模型架构下,多输入模态可提升嘈杂环境下自动语音识别(ASR)的准确性。通过在合成数据与真实场景数据上的实验发现:(1) 通常情况下,引入更多模态能提高ASR准确率,因各模态提供互补信息,但提升幅度取决于听觉噪声水平;(2) 同步模态(如唇动)在高噪声下更有用,而非同步模态(如图像上下文)在中等噪声下最有效;(3) 更高质量的视觉表征始终能提升ASR性能,凸显开发更强视觉编码器的重要性;(4) Mamba模型与Transformer在多模态收益趋势上表现出相似性;(5) 模态输入顺序及损失函数中的权重配置显著影响识别精度。这些发现既提供实用指导,也深化了对复杂条件下多模态语音识别的理解。
原文摘要 · Abstract (English)
Recent advances in multi-modal large language models (MLLMs) have opened new possibilities for unified modeling of speech, text, images, and other modalities. Building on our prior work, this paper examines the conditions and model architectures under which multiple input modalities can improve automatic speech recognition (ASR) accuracy in noisy environments. Through experiments on synthetic and real-world data, we find that (1) harnessing more modalities usually improves ASR accuracy, as each modality provides complementary information, but the improvement depends on the amount of auditory noise. (2) Synchronized modalities (e.g., lip movements) are more useful at high noise levels whereas unsynchronized modalities (e.g., image context) are most helpful at moderate noise levels. (3) Higher-quality visual representations consistently improve ASR accuracy, highlighting the importance of developing more powerful visual encoders. (4) Mamba exhibits similar trends regarding the benefits of multimodality as do Transformers. (5) The input order of modalities as well as their weights in the loss function can significantly impact accuracy. These findings both offer practical insights and help to deepen our understanding of multi-modal speech recognition under challenging conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。