arXiv:2502.17987cs.CLcs.AI2025-02被引 2

用多头注意力增强嵌入,提升稀缺语言情感分类效果

MAGE: Multi-Head Attention Guided Embeddings for Low Resource Sentiment Classification

  • 结合无语言依赖的数据增强与多头注意力加权嵌入
  • 在低资源班图语中显著提升情感分类准确率
  • 适合研究稀缺语言文本分类的学者与工程师

由于缺乏高质量的低资源班图语数据,文本分类及其他实际应用面临重大挑战。本文提出一种融合语言无关数据增强(LiDA)与基于多头注意力的加权嵌入的先进模型,通过选择性增强关键数据点来提升文本分类性能。该方法构建了适用于多种语言背景的鲁棒数据增强策略,使模型能有效处理班图语独特的句法和语义特征。不仅缓解了数据稀缺问题,也为未来低资源语言处理与分类任务奠定了基础。

原文摘要 · Abstract (English)

Due to the lack of quality data for low-resource Bantu languages, significant challenges are presented in text classification and other practical implementations. In this paper, we introduce an advanced model combining Language-Independent Data Augmentation (LiDA) with Multi-Head Attention based weighted embeddings to selectively enhance critical data points and improve text classification performance. This integration allows us to create robust data augmentation strategies that are effective across various linguistic contexts, ensuring that our model can handle the unique syntactic and semantic features of Bantu languages. This approach not only addresses the data scarcity issue but also sets a foundation for future research in low-resource language processing and classification tasks.

情感分类低资源语言多头注意力数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。