研究日语方言对语音与文本大模型的鲁棒性影响,发现两者表现相关且可提升。
Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language Models

- 以日语方言为测试场景,比较语音与文本大模型的方言理解能力。
- 语音模型的方言鲁棒性与对应文本模型表现高度相关。
- 加入方言数据训练和微调语音编码器能有效提升性能。
基于大语言模型(LLMs)的对话系统近年来发展迅速,但方言差异仍是主要挑战,尤其在处理语音输入时。将大语言模型与语音处理模块结合的语音语言模型(SLMs)在语音任务中展现出潜力,但其对方言的理解能力尚未充分研究。此外,基础文本模型的方言理解能力如何影响语音模型性能仍不明确。本研究以日语方言为案例,考察了文本与语音大模型的方言鲁棒性。我们定义鲁棒性为方言输入与标准输入下的性能比值,实现公平比较。实验表明,语音模型的鲁棒性与其对应的文本模型表现密切相关。进一步地,使用方言数据训练及微调语音编码器均能提升语音模型的鲁棒性。
原文摘要 · Abstract (English)
Dialogue systems based on large language models (LLMs) have advanced significantly in recent years. However, dialectal variation remains a major challenge, particularly for systems that process spoken input. LLM-based speech language models (SLMs), which integrate LLMs with speech processing components, show promise for spoken language tasks, yet their ability to comprehend dialects has not been sufficiently studied. Moreover, it remains unclear how the dialectal understanding of the base LLM affects SLM performance. This study investigates the dialectal robustness of both LLMs and SLMs using Japanese dialects as a test case. We define robustness as the ratio of performance on dialectal versus standard inputs, enabling fair comparisons. Our experiments show that SLM robustness correlates with that of their text-based counterparts. Furthermore, training with dialectal data and fine-tuning the speech encoder each improves robustness in SLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。