为假新闻视频生成可解释的自然语言说明,提升可信度判断透明度。
Multimodal Fake News Video Explanation: Dataset, Analysis and Evaluation
- 提出多模态关系图变换器,融合视频与标题信息进行推理
- 构建包含2672条数据的FakeVE数据集,覆盖四大类虚假特征
- 首次聚焦假新闻视频的解释生成,适合内容审核与AI可解释性研究
多模态假新闻视频难以理解,因其需综合分析视频与标题等多模态间的关联性与一致性。现有方法仅将其视为分类任务,缺乏对为何判定为假的解释。为此,我们提出新问题——假新闻视频解释(FNVE):给定含视频和标题的多模态新闻,目标是生成自然语言解释以揭示其虚假性。为此,我们构建了FakeVE数据集,包含2,672条可明确解释四类真实场景假新闻视频的样本。通过多角度探索性分析,我们深入理解了假新闻视频解释的特点。同时,提出基于多模态Transformer架构的多模态关系图变换器(MRGT)作为基准模型。实验证明,各基准模型在FakeVE上的表现令人信服,并提供了对不同模型解释生成差异的详细分析。
原文摘要 · Abstract (English)
Multimodal fake news videos are difficult to interpret because they require comprehensive consideration of the correlation and consistency between multiple modes. Existing methods deal with fake news videos as a classification problem, but it's not clear why news videos are identified as fake. Without proper explanation, the end user may not understand the underlying meaning of the falsehood. Therefore, we propose a new problem - Fake news video Explanation (FNVE) - given a multimodal news post containing a video and title, our goal is to generate natural language explanations to reveal the falsity of the news video. To that end, we developed FakeVE, a new dataset of 2,672 fake news video posts that can definitively explain four real-life fake news video aspects. In order to understand the characteristics of fake news video explanation, we conducted an exploratory analysis of FakeVE from different perspectives. In addition, we propose a Multimodal Relation Graph Transformer (MRGT) based on the architecture of multimodal Transformer to benchmark FakeVE. The empirical results show that the results of the various benchmarks (adopted by FakeVE) are convincing and provide a detailed analysis of the differences in explanation generation of the benchmark models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。