测试大模型对越南方言的鲁棒性,发现标准语表现好不代表方言下也可靠。
How Robust Are LLMs to Vietnamese Dialects?

- 构建首个系统化方言评测基准VialectBench,覆盖六类方言与四类任务。
- 方言输入导致平均性能下降2.82%,问答任务降幅最大,达6.17%。
- 中央方言组危害翻转率最高(6.54%),无模型完全具备方言不变性。
大型语言模型通常在标准书面越南语上进行评估,但日常交流中频繁使用保留语义但形式不同的地区方言。现有越南语方言研究多通过方言转标准语来处理问题,而非衡量模型在方言输入下的失效情况。为填补这一空白,我们首次系统评估了大模型在多个任务中对越南语方言变异的鲁棒性,量化了性能下降与失败模式。我们提出VialectBench(越南语方言评测基准),一个用于测试模型决策在六类越南语方言群体间是否稳定的受控基准。该基准包含400个标准越南语源实例及2,400条人工撰写的方言重写文本,涵盖情感识别(ER)、自然语言推理(NLI)、问答(QA)和多项选择题问答(MCQA)。固定参考模型评估显示,方言重写引发可测量的模型相对似然偏移,但长度与标准版本几乎相等。在十种指令微调模型上,方言输入使平均性能下降2.82%,且无一模型完全具备方言不变性。四项任务均受影响,问答任务平均下降最大。其中PNT3和PNT2方言组分别造成6.17%和4.73%的平均性能下降,而PNB方言组反而提升0.42%。中央方言组(PNT1-PNT4)在所有模型中产生最高的平均有害翻转率,达6.54%。结果表明,标准语上的优秀表现并不保证在意义保留的区域变体下行为可靠。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve meaning but differ in surface form. Existing Vietnamese dialect work largely addresses this issue through dialect-to-standard normalization instead of measuring how the model fails under Vietnamese dialectal inputs. To address this gap, we present the first systematic evaluation of LLM robustness to Vietnamese dialect variation across multiple tasks, quantifying performance degradation and failure patterns. We introduce VialectBench (Vietnamese Dialects Benchmarking), a controlled benchmark for testing whether model decisions remain stable across six Vietnamese dialect groups. VialectBench contains 400 Standard Vietnamese source instances and 2,400 human-written dialectal rewrites spanning emotion recognition (ER), natural language inference (NLI), question answering (QA), and multiple-choice question answering (MCQA). Dataset evaluation with a fixed reference language model shows that the dialectal rewrites induce a measurable model-relative likelihood shift while remaining nearly equal in length to their Standard counterparts. Across ten instruction-tuned models, dialectal inputs reduce average performance by 2.82%, and no evaluated model is fully dialect-invariant. All four tasks are affected, with QA showing the largest average degradation. Robustness also varies substantially across dialect groups: PNT3 and PNT2 cause the largest average performance drops, at 6.17% and 4.73%, respectively, whereas PNB slightly improves average performance by 0.42%. The Central dialect group (PNT1-PNT4) also yields the highest average harmful-flip rate across all models, at 6.54%. These findings show that strong performance on Standard Vietnamese does not guarantee reliable behavior under meaning-preserving regional variation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。