评测大模型在泰语方言上的表现,发现多数模型效果显著下降
Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation
- 构建涵盖北部、东北部、南部泰语方言的自动评估基准
- 五项任务中模型在方言上表现远低于标准泰语,仅少数闭源模型有较好表现
- 提出人类评估指南,兼顾方言准确性和生成流畅性,适合方言研究者参考
大语言模型在多种自然语言处理任务中表现出色,但在非主流语言及地方方言上的鲁棒性与一致性仍待探索。现有基准多聚焦于主流方言,忽视了地方方言文本的处理能力。本文引入一个覆盖泰语北部(兰纳)、东北部(依桑)和南部(丹布罗)方言的自动评估基准,评估大模型在摘要、问答、翻译、对话和食物相关任务中的表现。同时,提出针对泰语方言的人类评估指南与度量标准,以评估生成流畅性与方言特异性准确性。实验结果表明,与标准泰语相比,大模型在地方方言上的性能显著下降,仅有如 GPT-4o 与 Gemini2 等闭源模型展现出一定流畅性。
原文摘要 · Abstract (English)
Large language models show promising results in various NLP tasks. Despite these successes, the robustness and consistency of LLMs in underrepresented languages remain largely unexplored, especially concerning local dialects. Existing benchmarks also focus on main dialects, neglecting LLMs' ability on local dialect texts. In this paper, we introduce a Thai local dialect benchmark covering Northern (Lanna), Northeastern (Isan), and Southern (Dambro) Thai, evaluating LLMs on five NLP tasks: summarization, question answering, translation, conversation, and food-related tasks. Furthermore, we propose a human evaluation guideline and metric for Thai local dialects to assess generation fluency and dialect-specific accuracy. Results show that LLM performance declines significantly in local Thai dialects compared to standard Thai, with only proprietary models like GPT-4o and Gemini2 demonstrating some fluency
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。