arXiv:2607.04071cs.CLcs.AI2026-07中稿 · BRACIS 2026 - 36th…被引 1

首个专为葡萄牙语设计的句子编码评估基准,解决多语言模型在葡语表现不透明的问题。

Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders

论文配图:Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders
图 1 · 摘自论文原文
  • 构建14个葡语任务数据集组成的统一评估基准
  • 发现葡语表现高度依赖任务类型,无通用最优模型
  • 语言微调显著提升性能,尤其在对称监督任务上

葡萄牙语在全球使用广泛,但在文本嵌入评估中仍被忽视。现有模型常基于英语或多语言指标选择,其葡语实际表现不明。本文提出MTEB-PT,基于MMTEB子集构建,包含14个涵盖语义相似度、分类、检索和重排序的任务。我们用此基准统一评估17个开源与闭源模型。结果表明:葡语表现高度任务依赖,多语言排名无法可靠预测葡语表现;无单一模型在所有任务领先;长上下文能力强的模型在长输入任务(如检索)中优势明显。语言特定微调仍能提升性能,尤其在匹配训练数据的任务上。我们使用对比学习与马特约什卡表示学习对3个代表性模型进行微调,结果显示在语义相似度任务上增益最大,同时改善检索并保持维度截断下的竞争力。已发布MTEB-PT基准、微调模型及训练评估代码。

原文摘要 · Abstract (English)

Portuguese remains underrepresented in text embedding evaluation, despite being one of the most widely spoken languages in the world. As a result, embedding models are often selected based on English or multilingual metrics, while their effectiveness in Portuguese remains unclear. We present MTEB-PT, a Portuguese benchmark constructed from a subset of MMTEB, comprising 14 existing datasets across Semantic Textual Similarity (STS), classification, retrieval, and reranking. We use this benchmark to evaluate 17 open- and closed-source embedding models under a unified protocol. Our results show that Portuguese performance is strongly task-dependent: multilingual rankings do not reliably predict Portuguese-specific performance across task families, no single model dominates all settings, and models with stronger long-context capacity are particularly advantageous on longer-input tasks such as retrieval and reranking. The benchmark also shows that language-specific fine-tuning still improves model performance in Portuguese, especially on task types that match the adaptation data most closely. To examine this effect, we fine-tune three representative backbone models with Portuguese contrastive supervision and Matryoshka Representation Learning (MRL). These benchmark-informed baselines yield their strongest gains on STS, consistent with the predominantly symmetric supervision used during training, while also improving retrieval and remaining competitive under dimensional truncation. We release the MTEB-PT benchmark, the fine-tuned models, and the training and evaluation code.

句子编码葡萄牙语评估基准微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。