arXiv:2510.22356cs.CL2025-10

用翻译英文讽刺语料库,提升乌尔都语讽刺识别效果

Irony Detection in Urdu Text: A Comparative Study Using Machine Learning Models and Large Language Models

  • 将英文讽刺语料翻译为乌尔都语,构建本地化数据集
  • 大模型LLaMA 3(8B)达94.61%的F1分数,表现最佳
  • 为低资源语言提供可复用的讽刺检测方案

讽刺识别是自然语言处理中的难题,尤其在语法和文化背景差异大的语言中更为突出。本文通过将英文讽刺语料库翻译成乌尔都语,开展讽刺检测研究。我们评估了十种先进的机器学习算法,使用GloVe和Word2Vec词向量,并与传统方法对比。同时,对基于Transformer的大型模型BERT、RoBERTa、LLaMA 2(7B)、LLaMA 3(8B)和Mistral进行微调,以评估其在讽刺检测中的效果。在机器学习模型中,梯度提升(Gradient Boosting)表现最优,F1得分为89.18%;在变压器模型中,LLaMA 3(8B)取得最高成绩,F1得分为94.61%。结果表明,结合音译技术与现代NLP模型,可在乌尔都语这一历史低资源语言上实现稳健的讽刺识别。

原文摘要 · Abstract (English)

Ironic identification is a challenging task in Natural Language Processing, particularly when dealing with languages that differ in syntax and cultural context. In this work, we aim to detect irony in Urdu by translating an English Ironic Corpus into the Urdu language. We evaluate ten state-of-the-art machine learning algorithms using GloVe and Word2Vec embeddings, and compare their performance with classical methods. Additionally, we fine-tune advanced transformer-based models, including BERT, RoBERTa, LLaMA 2 (7B), LLaMA 3 (8B), and Mistral, to assess the effectiveness of large-scale models in irony detection. Among machine learning models, Gradient Boosting achieved the best performance with an F1-score of 89.18%. Among transformer-based models, LLaMA 3 (8B) achieved the highest performance with an F1-score of 94.61%. These results demonstrate that combining transliteration techniques with modern NLP models enables robust irony detection in Urdu, a historically low-resource language.

讽刺检测乌尔都语大模型低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。