arXiv:2509.06401cs.CL2025-09被引 2

测试大模型跨语言常识生成能力,发现英语表现远超其他语言。

Do LLMs exhibit the same commonsense capabilities across languages?

  • 构建多语言常识生成数据集MULTICOM,覆盖英西荷瓦伦西亚语。
  • 英语模型表现显著优于低资源语言,差距达20%以上。
  • 上下文提示对低资源语言有帮助,适合多语言研究者参考。

本文探究大型语言模型(LLMs)在多语言场景下的常识生成能力。为此,我们构建了MULTICOM——一个将COCOTEROS数据集扩展至四种语言(英语、西班牙语、荷兰语和瓦伦西亚语)的新基准。任务要求根据给定的三词组合生成符合常识的句子。我们评估了包括LLaMA、Qwen、Gemma、EuroLLM和Salamandra在内的多种开源模型。评估结合自动指标、基于LLM的评判方法(Prometheus和JudgeLM)以及人工标注。结果一致显示,英语表现最优,而低资源语言性能显著下降。尽管上下文支持效果不一,但对代表性不足的语言有一定提升作用。这些发现揭示了当前大模型在多语言常识生成方面的局限性。数据集已公开于https://huggingface.co/datasets/gplsi/MULTICOM。

原文摘要 · Abstract (English)

This paper explores the multilingual commonsense generation abilities of Large Language Models (LLMs). To facilitate this investigation, we introduce MULTICOM, a novel benchmark that extends the COCOTEROS dataset to four languages: English, Spanish, Dutch, and Valencian. The task involves generating a commonsensical sentence that includes a given triplet of words. We evaluate a range of open-source LLMs, including LLaMA, Qwen, Gemma, EuroLLM, and Salamandra, on this benchmark. Our evaluation combines automatic metrics, LLM-as-a-judge approaches (using Prometheus and JudgeLM), and human annotations. Results consistently show superior performance in English, with significantly lower performance in less-resourced languages. While contextual support yields mixed results, it tends to benefit underrepresented languages. These findings underscore the current limitations of LLMs in multilingual commonsense generation. The dataset is publicly available at https://huggingface.co/datasets/gplsi/MULTICOM.

大模型多语言常识推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。