用大模型分析方言语音,表现接近专业模型
Can LLM Agents Identify Spoken Dialects like a Linguist?
- 结合语音转录与语言学资源提升识别能力
- 提供人类专家和大模型基线对比
- 适合研究低资源方言的学者参考
由于缺乏标注的方言语音数据,大多数语言(包括瑞士德语)的音频方言分类任务极具挑战。本文探索大型语言模型(LLMs)作为智能体在理解方言方面的能力,并评估其性能是否可与HuBERT等模型相媲美。我们还提供了基于大模型和人类语言学家的基线。方法上,利用自动语音识别(ASR)系统生成的音素转录,并融合方言特征图、元音演变历史及语言规则等资源。结果表明,当引入语言学信息时,大模型的预测性能显著提升。人类基线显示,自动生成的转录有助于分类,但仍存在改进空间。
原文摘要 · Abstract (English)
Due to the scarcity of labeled dialectal speech, audio dialect classification is a challenging task for most languages, including Swiss German. In this work, we explore the ability of large language models (LLMs) as agents in understanding the dialects and whether they can show comparable performance to models such as HuBERT in dialect classification. In addition, we provide an LLM baseline and a human linguist one. Our approach uses phonetic transcriptions produced by ASR systems and combines them with linguistic resources such as dialect feature maps, vowel history, and rules. Our findings indicate that, when linguistic information is provided, the LLM predictions improve. The human baseline shows that automatically generated transcriptions can be beneficial for such classifications, but also presents opportunities for improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。