arXiv:2510.05026cs.CL2025-10ACL被引 3

用方言俗语测试大模型,发现多数模型不理解魁北克法语。

Idiom Understanding as a Tool to Measure the Dialect Gap

  • 用魁北克方言俗语构建新评测集,验证模型方言理解能力。
  • 111个大模型中65.77%在魁北克俗语上表现显著变差。
  • 适合评估模型对非主流方言的适应性,推动语言公平研究。

习语理解和方言理解均为自然语言处理中的成熟基准任务。本文提出将二者结合,以地区性习语作为方言理解的测试工具。为此,我们构建了三个针对魁北克法语的新基准数据集:QFrCoRE(含4,633条习语短语)、QFrCoRT(包含171个地区性习语词汇),以及一个用于法国本土表达的MFrCoE基准(含4,938个短语)。我们详细说明数据集构建方法,确保其他方言可复制。对111个大模型的实验显示,尽管模型在法国本土法语上表现良好,但65.77%的模型在魁北克习语上表现显著下降,仅9.0%偏向地区方言。结果证实这些基准能可靠量化方言差距,且母语标准语熟练度无法保证对地方方言的理解。

原文摘要 · Abstract (English)

The tasks of idiom understanding and dialect understanding are both well-established benchmarks in natural language processing. In this paper, we propose combining them, and using regional idioms as a test of dialect understanding. Towards this end, we propose three new benchmark datasets for the Quebec dialect of French: QFrCoRE, which contains 4,633 instances of idiomatic phrases, and QFrCoRT, which comprises 171 regional instances of idiomatic words, and a new benchmark for French Metropolitan expressions, MFrCoE, which comprises 4,938 phrases. We explain how to construct these corpora, so that our methodology can be replicated for other dialects. Our experiments with 111 LLMs reveal a critical disparity in dialectal competence: while models perform well on French Metropolitan, 65.77% of them perform significantly worse on Quebec idioms, with only 9.0% favoring the regional dialect. These results confirm that our benchmarks are a reliable tool for quantifying the dialect gap and that prestige-language proficiency does not guarantee regional dialect understanding.

方言理解习语识别大模型评估语言公平

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。