用多模态融合大模型精准识别短视频假新闻
VMID: A Multimodal Fusion LLM Framework for Detecting and Identifying Misinformation of Short Videos
- 融合视频图文信息生成统一文本,输入大模型分析
- 准确率达90.93%,优于基线模型的81.05%
- 适合关注短视频谣言治理的研究与从业者
短视频平台已成为新闻传播的重要渠道,但虚假信息也借其视觉吸引力和广泛传播力迅速扩散。现有检测方法多依赖单一模态或简单融合,难以应对短视频复杂的多层信息。本文提出一种基于多模态融合的大语言模型框架(VMID),通过多层次分析视频内容,整合不同模态特征生成统一文本描述,并输入大模型进行综合评估。实验表明,该方法在准确率上达90.93%,显著优于最佳基线模型SV-FEND的81.05%。案例研究进一步验证其在区分假新闻、辟谣内容与真实事件方面的可靠性与鲁棒性,适用于实际应用场景。
原文摘要 · Abstract (English)
Short video platforms have become important channels for news dissemination, offering a highly engaging and immediate way for users to access current events and share information. However, these platforms have also emerged as significant conduits for the rapid spread of misinformation, as fake news and rumors can leverage the visual appeal and wide reach of short videos to circulate extensively among audiences. Existing fake news detection methods mainly rely on single-modal information, such as text or images, or apply only basic fusion techniques, limiting their ability to handle the complex, multi-layered information inherent in short videos. To address these limitations, this paper presents a novel fake news detection method based on multimodal information, designed to identify misinformation through a multi-level analysis of video content. This approach effectively utilizes different modal representations to generate a unified textual description, which is then fed into a large language model for comprehensive evaluation. The proposed framework successfully integrates multimodal features within videos, significantly enhancing the accuracy and reliability of fake news detection. Experimental results demonstrate that the proposed approach outperforms existing models in terms of accuracy, robustness, and utilization of multimodal information, achieving an accuracy of 90.93%, which is significantly higher than the best baseline model (SV-FEND) at 81.05%. Furthermore, case studies provide additional evidence of the effectiveness of the approach in accurately distinguishing between fake news, debunking content, and real incidents, highlighting its reliability and robustness in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。