用少量波斯语数据实现高效情感分析,靠跨语言迁移与增量学习。
Cross-lingual Few-shot Learning for Persian Sentiment Analysis with Incremental Adaptation
- 利用多语言预训练模型,结合少样本与增量学习微调。
- mDeBERTa和XLM-RoBERTa在波斯语数据上达到96%准确率。
- 适合资源稀缺语言的情感分析研究者参考。
本研究探讨了在波斯语中使用少样本学习与增量学习进行跨语言情感分析的方法。目标是仅用有限数据实现波斯语情感分析,同时利用高资源语言的先验知识。实验采用三种预训练多语言模型(XLM-RoBERTa、mDeBERTa、DistilBERT),在来自X、Instagram、Digikala、Snappfood和Taaghche等多元来源的小规模波斯语数据上进行微调。多样化的数据来源使模型能够学习广泛语境下的表达。实验结果表明,mDeBERTa和XLM-RoBERTa在波斯语情感分析任务中均达到96%的准确率,证明了将少样本学习与增量学习结合多语言预训练模型的有效性。
原文摘要 · Abstract (English)
This research examines cross-lingual sentiment analysis using few-shot learning and incremental learning methods in Persian. The main objective is to develop a model capable of performing sentiment analysis in Persian using limited data, while getting prior knowledge from high-resource languages. To achieve this, three pre-trained multilingual models (XLM-RoBERTa, mDeBERTa, and DistilBERT) were employed, which were fine-tuned using few-shot and incremental learning approaches on small samples of Persian data from diverse sources, including X, Instagram, Digikala, Snappfood, and Taaghche. This variety enabled the models to learn from a broad range of contexts. Experimental results show that the mDeBERTa and XLM-RoBERTa achieved high performances, reaching 96% accuracy on Persian sentiment analysis. These findings highlight the effectiveness of combining few-shot learning and incremental learning with multilingual pre-trained models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。