首个专为巴西葡萄牙语设计的文本嵌入评测基准,打破依赖翻译数据的局限。
MTEB-BR: A Text Embedding Benchmark for Brazilian Portuguese

- 构建22个原生葡语任务,覆盖7类场景,杜绝翻译数据干扰。
- 93个模型实测显示:开源自托管模型可达顶级性能,无需商业API。
- 多语言排行榜预测葡语表现仅中等,凸显本地评测必要性。
葡萄牙语文本嵌入缺乏专用评测基准,现有评估依赖英文MS MARCO等翻译数据或覆盖稀疏的多语言数据集,且本土任务分散未整合。本文提出MTEB-BR,一个包含22个原生巴西葡萄牙语任务的评测基准,涵盖分类、多标签分类、成对分类、语义文本相似度、聚类、检索和重排序七类,所有数据均源自葡萄牙语原创或发现,不包含翻译内容。我们评估了93个模型(参数量从2300万到270亿),包括73个开放权重模型和20个封闭商业API。除排行榜外,还提供每项对比的统计层:任务级自助法置信区间、配对自助法显著性检验、基于项目反应理论的任务与实例级区分度分析,以及跨榜单相关性。三大发现突出:基准能清晰划分约十几个模型层级,但前六名在统计上无法排序;一个开源可自托管模型达到顶尖水平,说明优质葡语嵌入无需依赖商业接口;全球多语言排行榜对葡语排名预测能力有限(55个共享模型的斯皮尔曼相关系数为0.75,一模型在多语言榜第3,葡语榜第49),表明本地基准衡量的是多语言榜单未捕捉的特性。我们已公开所有任务、代码与公共排行榜,供从业者基于本土证据选择葡语嵌入模型。
原文摘要 · Abstract (English)
Text embeddings for Portuguese have no dedicated benchmark: evaluation rests on translated corpora such as English MS MARCO or on thin multilingual coverage, with native tasks scattered and unconsolidated. We introduce MTEB-BR, a benchmark of 22 native Brazilian-Portuguese tasks across seven categories (classification, multilabel classification, pair classification, semantic textual similarity, clustering, retrieval, and reranking), admitting only data created or found in Portuguese and excluding translations by construction. We evaluate 93 models spanning 23M to 27B parameters: 73 open-weight and 20 closed commercial APIs. Alongside the leaderboard we report a statistical layer for every headline comparison: per-task bootstrap confidence intervals, paired-bootstrap significance, a task- and instance-level discrimination analysis (how sharply each task separates models) adapted from Item Response Theory, and a cross-leaderboard correlation. Three findings stand out. The benchmark cleanly separates about a dozen tiers of models, though the top six are statistically too close to order. An openly licensed, self-hostable model reaches that leading tier, so strong Portuguese embedding quality does not require a commercial API. And a model's rank on the global multilingual leaderboard predicts its Portuguese rank only moderately (Spearman rho = 0.75 over 55 shared models; one model ranks 3rd there and 49th here), so a native benchmark measures something the multilingual boards do not. We release every task, our code, and a public leaderboard, so practitioners can choose Portuguese embedding models on native evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。