arXiv:2608.05510cs.CL2026-08

通过对比不同扰动方式,揭示了提升方言鲁棒性的机制差异。

Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness

论文配图:Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness
图 1 · 摘自论文原文
  • 系统比较六种扰动策略在九个方言任务中的表现
  • 字符噪声扰动显著提升零样本方言识别准确率
  • 相同效果可能源于不同适应机制,适合多语言场景选型

方言变体仍是多语言大模型面临的主要挑战。基于扰动的持续预训练(CPT)成为提升鲁棒性的有效方法,但现有研究多孤立评估单一扰动策略,缺乏对其作用机制的理解。本文对多语言大模型在德语、意大利语和阿拉伯语方言上的扰动式持续预训练进行了系统研究,比较了六种训练条件在九个方言任务上的表现。结果表明,基于扰动的CPT,尤其是字符噪声扰动,能稳定提升零样本方言鲁棒性,同时基本保持标准语种性能。更重要的是,即使下游表现相似,不同方法也表现出不同的鲁棒性机制:在语言模型适配、表征对齐和预测修复方面呈现各异模式。研究为理解合成表面变异如何增强鲁棒性提供了更完整的视角,并为多语言与方言场景下的CPT策略选择提供实践指导。

原文摘要 · Abstract (English)

Dialectal variation remains a major challenge for multilingual language models. Perturbation-based continued pre-training (CPT) has emerged as a promising approach to improving robustness, yet existing work largely evaluates individual perturbation strategies in isolation and provides limited insight into why they work. We present a systematic study of perturbation-based CPT for multilingual dialect robustness in LLMs, comparing six training conditions across nine German, Italian, and Arabic dialect tasks. Perturbation-based CPT, especially character-noised CPT, consistently improves zero-shot dialect robustness while largely preserving standard variety performance. More importantly, we show that methods with similar downstream performance induce distinct mechanisms of robustness, exhibiting different patterns of language model adaptation, representational alignment, and prediction repair. Our results provide a more complete understanding of how synthetic surface variation improves robustness and offer practical guidance for selecting CPT strategies in multilingual and dialectal settings.

多语言模型方言鲁棒性持续预训练机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。