构建首个罗马尼亚语多模态讽刺新闻数据集,提升讽刺检测效果
MuSaRoNews: A Multidomain, Multimodal Satire Dataset from Romanian News Articles
- 整合文本与视觉信息,构建多模态讽刺数据集
- 涵盖117,834篇真实与讽刺新闻,支持跨域分析
- 实验证明双模态融合显著提升检测性能
讽刺与假新闻虽目的不同(一为娱乐,一为误导),但均可能传播虚假信息。仅依赖文本难以捕捉新闻表面意义与实际含义之间的矛盾,视觉等其他信息源常提供关键线索。本文提出首个罗马尼亚语多模态讽刺新闻数据集MuSaRoNews,收集了117,834篇来自真实与讽刺新闻源的公开文章。实验表明,结合文本与视觉模态能有效提升讽刺检测性能。
原文摘要 · Abstract (English)
Satire and fake news can both contribute to the spread of false information, even though both have different purposes (one if for amusement, the other is to misinform). However, it is not enough to rely purely on text to detect the incongruity between the surface meaning and the actual meaning of the news articles, and, often, other sources of information (e.g., visual) provide an important clue for satire detection. This work introduces a multimodal corpus for satire detection in Romanian news articles named MuSaRoNews. Specifically, we gathered 117,834 public news articles from real and satirical news sources, composing the first multimodal corpus for satire detection in the Romanian language. We conducted experiments and showed that the use of both modalities improves performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。