用BERTurk在低资源土耳其语上训练情绪识别模型,分析反难民言论情感特征。
Emotion Recognition for Low-Resource Turkish: Fine-Tuning BERTurk on TREMO and Testing on Xenophobic Political Discourse
- 基于TREMO数据集微调BERTurk,构建专用于土耳其语的情绪识别模型。
- 在六类情绪分类上达到92.62%准确率,成功捕捉社交平台上的情感动态。
- 适用于舆情监控、危机管理等场景,推动小语种NLP应用落地。
社交媒体平台如X(前称推特)在塑造公共话语与社会规范中发挥关键作用。本研究聚焦土耳其社交媒体上的术语「Sessiz Istila」(沉默入侵),揭示叙利亚难民涌入背景下反难民情绪的上升趋势。利用BERTurk与TREMO数据集,我们开发了一种针对土耳其语的先进情绪识别模型(ERM),在幸福、恐惧、愤怒、悲伤、厌恶和惊讶六类情绪分类上实现了92.62%的准确率。通过将该模型应用于大规模X平台数据,研究揭示了土耳其语舆论中的情感细微差别,为计算社会科学提供了支持,推动了欠代表语言的情感分析发展,并深化了对全球数字话语及土耳其语独特语言挑战的理解。研究成果凸显本地化NLP工具的变革潜力,该ERM模型可在营销、公关与危机管理等领域实现实时情感分析,助力决策优化。这强调了关注区域与语言差异的研究的重要性。
原文摘要 · Abstract (English)
Social media platforms like X (formerly Twitter) play a crucial role in shaping public discourse and societal norms. This study examines the term Sessiz Istila (Silent Invasion) on Turkish social media, highlighting the rise of anti-refugee sentiment amidst the Syrian refugee influx. Using BERTurk and the TREMO dataset, we developed an advanced Emotion Recognition Model (ERM) tailored for Turkish, achieving 92.62% accuracy in categorizing emotions such as happiness, fear, anger, sadness, disgust, and surprise. By applying this model to large-scale X data, the study uncovers emotional nuances in Turkish discourse, contributing to computational social science by advancing sentiment analysis in underrepresented languages and enhancing our understanding of global digital discourse and the unique linguistic challenges of Turkish. The findings underscore the transformative potential of localized NLP tools, with our ERM model offering practical applications for real-time sentiment analysis in Turkish-language contexts. By addressing critical areas, including marketing, public relations, and crisis management, these models facilitate improved decision-making through timely and accurate sentiment tracking. This highlights the significance of advancing research that accounts for regional and linguistic nuances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。