arXiv:2511.17808cs.CL2025-11被引 6

首个大规模葡萄牙语大模型评估基准,揭示模型性能与投入的关系。

PoETa v2: Toward More Robust Evaluation of Large Language Models in Portuguese

  • 构建40+任务的葡萄牙语评测集PoETa v2,覆盖多种模型规模。
  • 测试超20个模型,发现计算资源投入越高,葡萄牙语表现越好。
  • 适合关注多语言大模型、特别是葡语研究的研究者使用。

大语言模型在不同语言和文化背景下的表现存在显著差异,亟需在多样语言中进行系统性评估。本文提出了迄今为止最全面的葡萄牙语大模型评估工作。基于我们新提出的PoETa v2基准——一个涵盖40多个葡萄牙语任务的综合性评测套件——我们评估了超过20个模型,覆盖广泛的训练规模和计算资源。研究揭示了计算投入与语言适配对葡萄牙语性能的影响,并分析了其与英语等效任务之间的性能差距。通过该基准与分析,PoETa v2为未来葡萄牙语语言建模与评估研究奠定了基础。基准数据集已公开于 https://github.com/PoETaV2/PoETaV2。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit significant variations in performance across linguistic and cultural contexts, underscoring the need for systematic evaluation in diverse languages. In this work, we present the most extensive evaluation of LLMs for the Portuguese language to date. Leveraging our newly introduced PoETa v2 benchmark -- a comprehensive suite of over 40 tasks in Portuguese -- we assess more than 20 models covering a broad spectrum of training scales and computational resources. Our study reveals how computational investment and language-specific adaptation impact performance in Portuguese, while also analyzing performance gaps in comparison to equivalent tasks in English. Through this benchmark and analysis, PoETa v2 lays the groundwork for future research on Portuguese language modeling and evaluation. The benchmark is available at https://github.com/PoETaV2/PoETaV2.

大模型评估葡萄牙语语言多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。