arXiv:2411.18099cs.CL2024-11被引 1

微调小型嵌入模型,显著提升尼泊尔语任务表现

Fine-Tuning Small Embeddings for Elevated Performance

  • 用不完整BERT模型在尼泊尔语上微调
  • 微调后性能远超原始基线,接近完整模型
  • 适合低资源语言的高效模型优化

上下文嵌入在自然语言处理任务中已取得顶尖效果,但其模型通常需要大量数据和巨大算力,对尼泊尔语等低资源语言构成挑战。本文采用一个仅含6个注意力头的不完整BERT模型,在尼泊尔语上进行预训练,并在未见数据上进行微调。通过内在与外在评估,结果与原模型基线及在尼泊尔语上完整预训练的基准模型(作为参照)对比。结果显示,尽管参照模型整体更优,但微调小模型仍显著优于原始基线,证明小模型经微调后可实现性能跃升。

原文摘要 · Abstract (English)

Contextual Embeddings have yielded state-of-the-art results in various natural language processing tasks. However, these embeddings are constrained by models requiring large amounts of data and huge computing power. This is an issue for low-resource languages like Nepali as the amount of data available over the internet is not always sufficient for the models. This work has taken an incomplete BERT model with six attention heads pretrained on Nepali language and finetuned it on previously unseen data. The obtained results from intrinsic and extrinsic evaluations have been compared to the results drawn from the original model baseline and a complete BERT model pretrained on Nepali language as the oracle. The results demonstrate that even though the oracle is better on average, finetuning the small embeddings drastically improves results compared to the original baseline.

嵌入模型低资源语言微调BERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。