用混合语料自适应XLM-RoBERTa,提升图鲁语混杂社交媒体中的希望言论检测
cantnlp@DravidianLangTech 2026: organic domain adaptation improves multi-class hope speech detection in Tulu
- 在真实图鲁语混杂社交文本上自适应微调XLM-RoBERTa模型
- 开发集上自适应模型准确率优于基线模型
- 适合关注低资源语言情感分析的研究者
本文报道了我们在第六届达罗毗荼语言语音、视觉与语言技术研讨会(DravidianLangTech-2026)上针对图鲁语代码混杂语境下的希望言论检测任务的系统与结果。我们基于XLM-RoBERTa构建了文本分类模型,用于识别图鲁语社交媒体评论中的希望言论。通过对比在自然收集的图鲁语混杂社交文本上进行自适应训练的模型与基线模型的表现,发现前者在开发集上取得更优效果。尽管提交系统在官方测试集上表现较为一般,但结果表明,在包含代码混杂与多书写系统变异的真实语料上进一步适配XLM-RoBERTa,有助于提升图鲁语混杂语境下的希望言论检测性能。
原文摘要 · Abstract (English)
This paper presents our systems and results for the Hope Speech Detection in Code-Mixed Tulu Language shared task at the Sixth Workshop on Speech, Vision, and Language Technologies for Dravidian Languages (DravidianLangTech-2026). We trained an XLM-RoBERTa-based text classification system for detecting hope speech in code-mixed Tulu social media comments. We compared this organically adapted hope speech detection model with our baseline model. On the development set, the organically adapted model outperformed the baseline system. While our submitted systems performed more modestly on the official test set, these results suggest that further adapting XLM-RoBERTa on organically collected Tulu social media text containing code-mixed and mixed-script variation can improve hope speech detection in code-mixed Tulu.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。