用纯标题检测罗马尼亚新闻讽刺语调,效果优于传统方法。
SaRoHead: Detecting Satire in a Multi-Domain Romanian News Headline Dataset
- 仅用标题检测讽刺,不依赖正文内容。
- 双向Transformer模型表现最佳,尤其结合元学习Reptile时。
- 适合关注多领域讽刺文本检测的研究者和媒体分析人员。
新闻标题的首要目标是用最少的字数概括事件。根据媒体风格,标题可客观陈述事实,也可通过讽刺、反语和夸张等手法提升传播力。因此,标题需反映主文的讽刺基调。当前针对罗马尼亚语的讽刺检测方法通常结合文章正文与标题。本文认为标题本身已是内容的简要概括,故聚焦于仅凭标题识别讽刺语气,测试了从传统机器学习到深度学习模型的多种基线方法。实验表明,双向Transformer模型在性能上超越标准机器学习方法及大语言模型(LLMs),尤其在采用元学习算法Reptile时表现更优。
原文摘要 · Abstract (English)
The primary goal of a news headline is to summarize an event in as few words as possible. Depending on the media outlet, a headline can serve as a means to objectively deliver a summary or improve its visibility. For the latter, specific publications may employ stylistic approaches that incorporate the use of sarcasm, irony, and exaggeration, key elements of a satirical approach. As such, even the headline must reflect the tone of the satirical main content. Current approaches for the Romanian language tend to detect the non-conventional tone (i.e., satire and clickbait) of the news content by combining both the main article and the headline. Because we consider a headline to be merely a brief summary of the main article, we investigate in this paper the presence of satirical tone in headlines alone, testing multiple baselines ranging from standard machine learning algorithms to deep learning models. Our experiments show that Bidirectional Transformer models outperform both standard machine-learning approaches and Large Language Models (LLMs), particularly when the meta-learning Reptile approach is employed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。