arXiv:2510.05046cs.CL2025-10被引 2

构建首个覆盖23项任务的法语理解评测基准,揭示大模型在法语上的性能差距。

COLE: a Comprehensive Benchmark for French Language Understanding Evaluation

  • 构建23个多样化任务,涵盖情感分析、推理等法语特有语言现象。
  • 94个大模型对比显示闭源模型显著优于开源模型,零样本问答等任务仍具挑战。
  • 专为法语设计,适合法语NLP研究者与多语言模型开发者参考。

为应对法语自然语言理解(NLU)评估的不足,我们提出COLE,一个包含23个多样化任务的综合性评测基准,覆盖情感分析、改写检测、语法判断和推理等广泛能力,尤其关注法语特有的语言现象。我们对94个大型语言模型进行了评测,全面分析了当前法语NLU的水平。结果表明,闭源模型与开源模型之间存在显著性能差距,当前模型在零样本抽取式问答、细粒度词义消歧以及区域语言变体理解方面仍面临关键挑战。COLE作为公开资源发布,旨在推动法语语言建模的进一步发展。

原文摘要 · Abstract (English)

To address the need for a more comprehensive evaluation of French Natural Language Understanding (NLU), we introduce COLE, a new benchmark composed of 23 diverse task covering a broad range of NLU capabilities, including sentiment analysis, paraphrase detection, grammatical judgment, and reasoning, with a particular focus on linguistic phenomena relevant to the French language. We benchmark 94 large language models (LLM), providing an extensive analysis of the current state of French NLU. Our results highlight a significant performance gap between closed- and open-weights models and identify key challenging frontiers for current LLMs, such as zero-shot extractive question-answering (QA), fine-grained word sense disambiguation, and understanding of regional language variations. We release COLE as a public resource to foster further progress in French language modelling.

法语NLP评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。