对比推特与TikTok,发现假消息引发不同情绪,音频特征可提升识别效果。
Divergent Emotional Patterns in Disinformation on Social Media? An Analysis of Tweets and TikToks about the DANA in Valencia
- 分析推特与TikTok内容情绪差异,揭示假消息在不同平台引发不同情感反应。
- 假消息多用否定、感知词和个人故事,可信内容语言更正式客观。
- 音频中情感化音调与音乐增强传播力,融合音频特征可显著提升检测准确率。
本研究分析2024年10月29日西班牙瓦伦西亚因极端降雨引发洪灾期间,社交媒体上关于DANA(高海拔孤立低压)事件的虚假信息传播。构建了包含650条TikTok和X平台帖子的新型数据集,并通过人工标注区分虚假与可信内容。基于GPT-4o的少样本标注方法与人工标签达成较高一致性(Cohen's kappa=0.684)。情感分析显示,推特上的假消息主要关联悲伤与恐惧,而TikTok则关联愤怒与厌恶。使用LIWC词典的语义分析表明,可信内容语言更严谨客观,假消息则依赖否定句、感知动词和个人叙事增强可信度。音频分析发现,可信内容音频以明亮音色和机械式单调表达为主,假消息则运用音调变化、情感深度及操纵性配乐提升互动性。模型方面,SVM+TF-IDF在小样本下表现最佳;将音频特征引入roberta-large-bne后,准确率与F1分数均优于纯文本模型与SVM。GPT-4o少样本方法也表现良好,显示出大语言模型在自动化识别中的潜力。结果表明,结合文本与音频特征对多模态平台的虚假信息检测至关重要。
原文摘要 · Abstract (English)
This study investigates the dissemination of disinformation on social media platforms during the DANA event (DANA is a Spanish acronym for Depresion Aislada en Niveles Altos, translating to high-altitude isolated depression) that resulted in extremely heavy rainfall and devastating floods in Valencia, Spain, on October 29, 2024. We created a novel dataset of 650 TikTok and X posts, which was manually annotated to differentiate between disinformation and trustworthy content. Additionally, a Few-Shot annotation approach with GPT-4o achieved substantial agreement (Cohen's kappa of 0.684) with manual labels. Emotion analysis revealed that disinformation on X is mainly associated with increased sadness and fear, while on TikTok, it correlates with higher levels of anger and disgust. Linguistic analysis using the LIWC dictionary showed that trustworthy content utilizes more articulate and factual language, whereas disinformation employs negations, perceptual words, and personal anecdotes to appear credible. Audio analysis of TikTok posts highlighted distinct patterns: trustworthy audios featured brighter tones and robotic or monotone narration, promoting clarity and credibility, while disinformation audios leveraged tonal variation, emotional depth, and manipulative musical elements to amplify engagement. In detection models, SVM+TF-IDF achieved the highest F1-Score, excelling with limited data. Incorporating audio features into roberta-large-bne improved both Accuracy and F1-Score, surpassing its text-only counterpart and SVM in Accuracy. GPT-4o Few-Shot also performed well, showcasing the potential of large language models for automated disinformation detection. These findings demonstrate the importance of leveraging both textual and audio features for improved disinformation detection on multimodal platforms like TikTok.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。