解决泰米尔-英语混用文本情感分析难题,提升多语言模型表现
Advancing Sentiment Analysis in Tamil-English Code-Mixed Texts: Challenges and Transformer-Based Solutions
- 采用XLM-RoBERTa等多语言Transformer模型处理混用文本
- 现有数据集规模小且标注不全,影响模型性能上限
- 适合多语言自然语言处理研究者与跨语言情感分析应用
针对泰米尔-英语混用文本的情感分析任务,研究探索了基于先进Transformer模型的解决方案。针对语法不一致、拼写差异和语音歧义等挑战进行了分析。现有数据集局限性及标注空白被重点考察,强调需构建更大更丰富的语料库。在低资源环境下评估了XLM-RoBERTa、mT5、IndicBERT和RemBERT等架构的性能,结果显示特定模型在多语言情感分类中表现优异。研究指出,未来需通过数据增强、语音标准化和混合建模进一步提升准确率,并提出若干发展方向。
原文摘要 · Abstract (English)
The sentiment analysis task in Tamil-English code-mixed texts has been explored using advanced transformer-based models. Challenges from grammatical inconsistencies, orthographic variations, and phonetic ambiguities have been addressed. The limitations of existing datasets and annotation gaps have been examined, emphasizing the need for larger and more diverse corpora. Transformer architectures, including XLM-RoBERTa, mT5, IndicBERT, and RemBERT, have been evaluated in low-resource, code-mixed environments. Performance metrics have been analyzed, highlighting the effectiveness of specific models in handling multilingual sentiment classification. The findings suggest that further advancements in data augmentation, phonetic normalization, and hybrid modeling approaches are required to enhance accuracy. Future research directions for improving sentiment analysis in code-mixed texts have been proposed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。