情绪影响语言选择:积极内容用更多英语,切换更频繁。
EnTaCs: Analyzing the Relationship Between Sentiment and Language Choice in English-Tamil Code-Switching
- 用XLM-RoBERTa分析3.5万条英泰混杂评论,量化每句英语比例和切换频率。
- 积极语句英语占比34.3%,远高于消极语句的24.8%;中性情绪切换最频繁。
- 发现情感通过社会地位与身份认同影响语言混合行为,适合研究多语者交流者。
本文研究英语-泰米尔语混用文本中语句情感与语言选择的关系,采用机器学习与统计建模方法。基于DravidianCodeMix数据集中的35,650条罗马化YouTube评论,使用微调后的XLM-RoBERTa模型进行细粒度语言识别,获得每条语句的英语占比与语言切换频率。线性回归分析显示,积极语句的英语占比显著更高(34.3%),而消极语句仅为24.8%;在控制语句长度后,混合情感语句的切换频率最高。结果支持情感内容会通过社会语言学中的地位与身份关联,影响多语言混用中的语言选择。
原文摘要 · Abstract (English)
This paper investigates the relationship between utterance sentiment and language choice in English-Tamil code-switched text, using methods from machine learning and statistical modelling. We apply a fine-tuned XLM-RoBERTa model for token-level language identification on 35,650 romanized YouTube comments from the DravidianCodeMix dataset, producing per-utterance measurements of English proportion and language switch frequency. Linear regression analysis reveals that positive utterances exhibit significantly greater English proportion (34.3%) than negative utterances (24.8%), and mixed-sentiment utterances show the highest language switch frequency when controlling for utterance length. These findings support the hypothesis that emotional content demonstrably influences language choice in multilingual code-switching settings, due to socio-linguistic associations of prestige and identity with embedded and matrix languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。