分析五类新闻中可信与虚假信息的语言差异,发现需按领域适配检测方法。
Variation between Credible and Non-Credible News Across Topics
- 对比经济、娱乐、健康、科学、体育五类新闻的语言风格差异
- 不同领域中可信与虚假新闻的语义特征存在显著区别
- 建议按主题定制检测模型,提升实际应用效果
虚假新闻持续损害现代新闻业与政治信任。尽管研究不断,结果仍不一致。以往研究多聚焦于真假新闻区分或子类型(如宣传、讽刺、错误信息)的识别。本文从语言学与文体角度分析虚假新闻,关注其在不同新闻主题间的差异。基于话语与语言学中的欺骗检测相关工作,分析经济、娱乐、健康、科学和体育五个具体主题。结果显示,各类别中可信与虚假新闻的语言特征存在显著差异,强调分类任务需考虑主题相关的风格与语言特性,以提升真实场景下的表现。
原文摘要 · Abstract (English)
'Fake News' continues to undermine trust in modern journalism and politics. Despite continued efforts to study fake news, results have been conflicting. Previous attempts to analyse and combat fake news have largely focused on distinguishing fake news from truth, or differentiating between its various sub-types (such as propaganda, satire, misinformation, etc.) This paper conducts a linguistic and stylistic analysis of fake news, focusing on variation between various news topics. It builds on related work identifying features from discourse and linguistics in deception detection by analysing five distinct news topics: Economy, Entertainment, Health, Science, and Sports. The results emphasize that linguistic features vary between credible and deceptive news in each domain and highlight the importance of adapting classification tasks to accommodate variety-based stylistic and linguistic differences in order to achieve better real-world performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。