无需标注数据,通过自生成标签提升语音大模型在特定场景的表现。
Self-Improvement for Audio Large Language Model using Unlabeled Speech
- 利用大模型解码信息评估伪标签质量,结合强化学习实现领域自适应。
- 在多个ASR、SQA和S2TT数据集上,WER与BLEU指标均显著优于基线。
- 仅需少量数据即可优化模型,适合真实场景部署。
近期语音大语言模型迅速发展,在多种语音任务中展现出强大泛化能力。然而,由于语音信号固有的复杂性,这些模型在特定目标领域性能仍会下降。为解决此问题,本文提出一种无需标注数据的自改进方法SI-SDA,利用大模型解码过程中的信息评估生成伪标签的质量,并基于强化学习进行领域适配。实验表明,该方法在多个公开的自动语音识别(ASR)、口语问答(SQA)和语音到文本翻译(S2TT)数据集上,持续且显著提升音频大模型性能,各项指标(如WER和BLEU)均优于现有基线。此外,本方法表现出高数据效率,具备良好的实际应用潜力。
原文摘要 · Abstract (English)
Recent audio LLMs have emerged rapidly, demonstrating strong generalization across various speech tasks. However, given the inherent complexity of speech signals, these models inevitably suffer from performance degradation in specific target domains. To address this, we focus on enhancing audio LLMs in target domains without any labeled data. We propose a self-improvement method called SI-SDA, leveraging the information embedded in large-model decoding to evaluate the quality of generated pseudo labels and then perform domain adaptation based on reinforcement learning optimization. Experimental results show that our method consistently and significantly improves audio LLM performance, outperforming existing baselines in WER and BLEU across multiple public datasets of automatic speech recognition (ASR), spoken question-answering (SQA), and speech-to-text translation (S2TT). Furthermore, our approach exhibits high data efficiency, underscoring its potential for real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。