首个罗马尼亚语讽刺句级检测数据集,助力自然语言理解研究
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
- 构建13,873条罗马尼亚语新闻句子的讽刺标注数据集
- 大模型在零样本和微调设置下均表现有限,揭示现有模型不足
- 适合关注多语言讽刺识别与LLM能力边界的学者
讽刺、反语和挖苦常用于表达幽默与批评,而非误导;但有时可能被误认为事实报道,类似假新闻。这些手法可应用于更细粒度层面,使讽刺内容融入新闻文章。本文提出首个面向罗马尼亚语新闻文章的句级讽刺检测数据集——SeLeRoSa,包含13,873条人工标注的句子,覆盖社会议题、信息技术、科学和电影等多个领域。随着大规模语言模型(LLMs)在自然语言处理中的兴起与进展,其在零样本设置下已展现出强大任务处理能力。我们评估了多种基于LLM的基线模型在零样本与微调设置下的表现,以及传统Transformer模型。结果表明,当前模型在句级讽刺检测任务中仍存在明显局限,为未来研究指明新方向。
原文摘要 · Abstract (English)
Satire, irony, and sarcasm are techniques typically used to express humor and critique, rather than deceive; however, they can occasionally be mistaken for factual reporting, akin to fake news. These techniques can be applied at a more granular level, allowing satirical information to be incorporated into news articles. In this paper, we introduce the first sentence-level dataset for Romanian satire detection for news articles, called SeLeRoSa. The dataset comprises 13,873 manually annotated sentences spanning various domains, including social issues, IT, science, and movies. With the rise and recent progress of large language models (LLMs) in the natural language processing literature, LLMs have demonstrated enhanced capabilities to tackle various tasks in zero-shot settings. We evaluate multiple baseline models based on LLMs in both zero-shot and fine-tuning settings, as well as baseline transformer-based models. Our findings reveal the current limitations of these models in the sentence-level satire detection task, paving the way for new research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。