测试10种东南亚语言嵌入模型表现,发现无模型通用优,任务间差异大。
SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
- 构建覆盖10种东南亚语言的大规模基准SEA-BED,评估多任务表现
- 各模型在不同语言任务中表现不一,无统一最优模型
- 揭示语言-任务组合性能高度不均,需全面评估语义表征
多语言文本嵌入通常假设在与视角无关的语义空间中编码意义,从而在不同任务和语言间保持稳定的相似性判断。但我们的结果表明这一假设在实践中并不成立。本文提出SEA-BED,一个涵盖10种东南亚(SEA)语言及多种嵌入任务的大型基准,旨在系统检验嵌入性能在任务、语言及语言-任务组合间的差异。大规模评估显示,无单一模型在所有东南亚语言中表现一致;同一语言内任务难度差异显著,且某任务成功无法可靠泛化至其他任务。语言-任务分析进一步揭示了高度非均匀的性能分布,性能随语言-任务组合变化剧烈。这些发现呼吁更全面地测量性能,以揭示语义表征中的不一致性。基于此,我们为未来模型开发提供数据、算法与架构层面的启示。
原文摘要 · Abstract (English)
Multilingual text embeddings are often assumed to encode meaning in a perspective-independent semantic space, yielding stable similarity judgments across tasks and languages. Our results show that this assumption does not hold in practice. We introduce SEA-BED, a large-scale benchmark covering 10 Southeast Asian (SEA) languages and diverse embedding tasks, designed to systematically examine how embedding performance varies across tasks, languages, and language-task combinations. Across extensive evaluations, we observe that no single model performs uniformly well across SEA languages; task difficulty differs markedly within languages, and success on one task does not reliably generalize to others. Language-task analyses further reveal highly non-uniform performance landscapes, where performance varies across different language-task combinations. These findings call for closer attention to performance measurements that provide an expansive view across languages and tasks to uncover inconsistencies in semantic representation. Based on these observations, we provide insights for future model development, including data, algorithmic, and architectural considerations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。