arXiv:2606.03027cs.CL2026-06

开源可复现的东南亚语言文本嵌入模型,提升本地化NLP鲁棒性。

SEA-Embedding: Open and Reproducible Text Embeddings for Southeast Asia

论文配图:SEA-Embedding: Open and Reproducible Text Embeddings for Southeast Asia
图 1 · 摘自论文原文
  • 基于公开数据构建全开源嵌入管道,支持复现
  • 在SEA-BED上达到当前最佳性能
  • 适合研究东南亚语言NLP的学者与开发者

文本嵌入是众多下游应用的基础,其鲁棒性对实际NLP应用至关重要。然而,当前多数顶尖嵌入模型因依赖封闭或未公开的训练数据而难以复现,且对东南亚语言的适应性不足。本文提出SEA-Embedding,一个完全开源、可复现的东南亚语言文本嵌入方案,仅使用公开数据训练,并系统研究了鲁棒嵌入设计的三大核心因素:数据构成、训练目标与基础编码器初始化。该模型在SEA-BED基准上取得领先性能,同时支持对区域语言嵌入的可复现分析。

原文摘要 · Abstract (English)

Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP. However, most recent state-of-the-art embedding models are not reproducible because they rely on closed or undisclosed training data, and they remain insufficiently robust for Southeast Asian languages. We present SEA-Embedding, a fully open and reproducible text-embedding pipeline for Southeast Asian languages trained only on publicly available data, and use it to study three core factors of robust embedding design: data composition, training objective, and base encoder initialization. SEA-Embedding achieves state-of-the-art results on SEA-BED while enabling systematic and reproducible analysis of robust text embeddings for the region.

文本嵌入东南亚语言开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。