arXiv:2411.15270cs.CLcs.LG2024-11中稿 · ACAI 2024被引 5

用跨语言蒸馏让低资源语言邦加尔语实现高效句向量表示

BanglaEmbed: Efficient Sentence Embedding Models for a Low-Resource Language Using Cross-Lingual Distillation Techniques

  • 通过英文优秀模型蒸馏知识,构建邦加尔语轻量句向量模型
  • 在相似度、同义检测等任务上超越现有邦加尔语模型
  • 模型小速度快,适合计算资源有限的部署场景

句级嵌入对自然语言理解任务至关重要。尽管英语等高资源语言已有大量研究,但像孟加拉语(约2.3亿人使用)这样的低资源语言仍被严重忽视。本文提出两种轻量级孟加拉语句向量模型,采用新颖的跨语言知识蒸馏方法,从高性能英文句向量模型中迁移知识。所提模型在多个下游任务中评估,包括句子相似度(STS)、同义检测和孟加拉语仇恨言论识别,均显著优于现有孟加拉语句向量模型。此外,其轻量化架构和更短推理时间使其特别适用于资源受限环境中的实际自然语言处理应用。

原文摘要 · Abstract (English)

Sentence-level embedding is essential for various tasks that require understanding natural language. Many studies have explored such embeddings for high-resource languages like English. However, low-resource languages like Bengali (a language spoken by almost two hundred and thirty million people) are still under-explored. This work introduces two lightweight sentence transformers for the Bangla language, leveraging a novel cross-lingual knowledge distillation approach. This method distills knowledge from a pre-trained, high-performing English sentence transformer. Proposed models are evaluated across multiple downstream tasks, including paraphrase detection, semantic textual similarity (STS), and Bangla hate speech detection. The new method consistently outperformed existing Bangla sentence transformers. Moreover, the lightweight architecture and shorter inference time make the models highly suitable for deployment in resource-constrained environments, making them valuable for practical NLP applications in low-resource languages.

句向量低资源语言知识蒸馏孟加拉语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。