用主动学习与聚类提升低资源语言情感分析效果
Enhancing BERT Fine-Tuning for Sentiment Analysis in Lower-Resourced Languages
- 结合主动学习与聚类的动态数据筛选策略
- 标注量减少30%且F1分数最高提升4分
- 适合数据稀缺但需稳定模型的场景
低资源语言因数据有限,其语言模型性能较弱。由于预训练计算成本高,更实际的做法是优化微调阶段。本文研究在有限数据下,通过结合主动学习(AL)与结构化数据选择策略(称为‘主动学习调度器’)来提升微调效果。我们引入聚类机制,构建集成微调流程,系统融合主动学习、聚类和动态数据选择调度器。在斯洛伐克语、马耳他语、冰岛语和土耳其语上的实验表明,微调阶段结合聚类与主动学习调度可同时实现最高30%的标注量节省和最高4个F1分数提升,且微调过程更稳定。
原文摘要 · Abstract (English)
Limited data for low-resource languages typically yield weaker language models (LMs). Since pre-training is compute-intensive, it is more pragmatic to target improvements during fine-tuning. In this work, we examine the use of Active Learning (AL) methods augmented by structured data selection strategies which we term 'Active Learning schedulers', to boost the fine-tuning process with a limited amount of training data. We connect the AL to data clustering and propose an integrated fine-tuning pipeline that systematically combines AL, clustering, and dynamic data selection schedulers to enhance model's performance. Experiments in the Slovak, Maltese, Icelandic and Turkish languages show that the use of clustering during the fine-tuning phase together with AL scheduling can simultaneously produce annotation savings up to 30% and performance improvements up to four F1 score points, while also providing better fine-tuning stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。