arXiv:2503.04786cs.CLcs.SI2025-03

分析虚假信息语言特征随时间演变,揭示其情绪与来源变化规律。

Analyzing the temporal dynamics of linguistic features contained in misinformation

  • 用NLP分析2010-2024年PolitiFact数据,分五阶段追踪语言特征动态。
  • 虚假信息语气更负面,近年多来自社交媒体和网络论坛,含大量负面情绪内容。
  • 虚假信息常涉及总统候选人等人物实体,准确信息多含百分比、日期等数字实体。

虚假信息的传播可能对个人和社会造成负面影响。为减轻其对公众信念的影响,已有算法通过内容准确性和来源可靠性提供标注。由于算法用于评估信息准确性的语言特征会随时间变化,理解其时序动态至关重要。本研究利用自然语言处理技术,分析2010至2024年间PolitiFact的陈述,量化不同五年周期内虚假信息的来源与语言特征变化。结果显示,陈述情感值显著下降,整体语气趋于负面;与真实信息相比,虚假信息的情感值更低。进一步分析发现,近期陈述主要来自社交媒体、博客及病毒式图片等数字平台,其中虚假信息占比高且情绪消极。而早期陈述多源自政治人物等个体来源,准确性相对均衡,语调中性或积极。命名实体识别显示,总统现任者与候选人更常出现在虚假信息中,美国各州则更多见于真实信息。此外,虚假信息中人物与组织类实体更常见,真实陈述则更常包含百分比、日期等数值实体。

原文摘要 · Abstract (English)

Consumption of misinformation can lead to negative consequences that impact the individual and society. To help mitigate the influence of misinformation on human beliefs, algorithmic labels providing context about content accuracy and source reliability have been developed. Since the linguistic features used by algorithms to estimate information accuracy can change across time, it is important to understand their temporal dynamics. As a result, this study uses natural language processing to analyze PolitiFact statements spanning between 2010 and 2024 to quantify how the sources and linguistic features of misinformation change between five-year time periods. The results show that statement sentiment has decreased significantly over time, reflecting a generally more negative tone in PolitiFact statements. Moreover, statements associated with misinformation realize significantly lower sentiment than accurate information. Additional analysis shows that recent time periods are dominated by sources from online social networks and other digital forums, such as blogs and viral images, that contain high levels of misinformation containing negative sentiment. In contrast, most statements during early time periods are attributed to individual sources (i.e., politicians) that are relatively balanced in accuracy ratings and contain statements with neutral or positive sentiment. Named-entity recognition was used to identify that presidential incumbents and candidates are relatively more prevalent in statements containing misinformation, while US states tend to be present in accurate information. Finally, entity labels associated with people and organizations are more common in misinformation, while accurate statements are more likely to contain numeric entity labels, such as percentages and dates.

虚假信息语言特征时间演化NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。