arXiv:2410.12704cs.LGcs.CL2024-10被引 3

用翻译+大模型构建斯洛文尼亚语讽刺检测数据集

Sarcasm Detection in a Less-Resourced Language

  • 用翻译模型和生成式大模型构建斯洛文尼亚语讽刺数据集
  • 大模型比小模型表现更好,集成方法微幅提升效果
  • 为低资源语言提供可复用的讽刺识别方案

自然语言处理中的讽刺检测旨在判断话语是否具有讽刺意味,与情感分析密切相关,常表现为表层情感反转。由于讽刺高度依赖上下文及非语言线索,任务极具挑战性。现有研究多集中于英语等高资源语言。为构建低资源语言(如斯洛文尼亚语)的讽刺检测数据集,本文采用中等规模机器翻译专用Transformer模型与超大规模生成式语言模型,探索了翻译数据集的可行性及预训练模型大小对讽刺识别能力的影响。通过训练多种检测模型集成系统并评估性能,结果表明:更大模型普遍优于小模型,集成策略能轻微提升性能。最佳集成模型在斯洛文尼亚语上达到F₁=0.765,接近原始语言标注者间一致性水平。

原文摘要 · Abstract (English)

The sarcasm detection task in natural language processing tries to classify whether an utterance is sarcastic or not. It is related to sentiment analysis since it often inverts surface sentiment. Because sarcastic sentences are highly dependent on context, and they are often accompanied by various non-verbal cues, the task is challenging. Most of related work focuses on high-resourced languages like English. To build a sarcasm detection dataset for a less-resourced language, such as Slovenian, we leverage two modern techniques: a machine translation specific medium-size transformer model, and a very large generative language model. We explore the viability of translated datasets and how the size of a pretrained transformer affects its ability to detect sarcasm. We train ensembles of detection models and evaluate models' performance. The results show that larger models generally outperform smaller ones and that ensembling can slightly improve sarcasm detection performance. Our best ensemble approach achieves an $\text{F}_1$-score of 0.765 which is close to annotators' agreement in the source language.

讽刺检测低资源语言机器翻译模型集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。