新造假视频泛滥,这篇论文建了首个纯合成假新闻数据集并提出新检测框架。
From Cheap Fakes to Pure Synthesis: Addressing the New Era of T2V Fake News Videos

- 将假新闻视频检测升级为三类分类任务,区分真实、廉价伪造和纯合成伪造。
- 构建首个纯合成假新闻视频数据集,包含两类伪造事件,避免模型走捷径。
- 提出基于推理的检测框架,结合语义逻辑与生成痕迹,准确率领先12.2个百分点。
近期文本到视频(T2V)生成模型可从零生成假新闻视频,使威胁超越由现有影像拼接的廉价伪造。此类视频能精准匹配虚构叙事,形成模态对齐陷阱,现有检测器难以应对。现有数据集缺乏纯合成假新闻视频。尽管直接用假新闻描述提示T2V模型可生成完全对齐样本,但会使假新闻视频检测(FNVD)退化为单模态捷径,并导致语义-视觉退化。为此,我们提出将T2V-FNVD定义为新型三分类任务,包含真实、廉价伪造和纯合成伪造三类标签,并构建首个纯合成假新闻视频数据集(PS-FNVD)。该数据集包含两类伪造:与虚构叙事对齐的虚构事件(类型1)和具有虚假视觉来源的真实事件(类型2),有效防止模型依赖单模态捷径。此外,我们提出推理引导的T2V-FNVD(R-T2V)框架。通过条件推理生成与监督微调训练,R-T2V融合高层语义逻辑与低层物理生成痕迹,预测三元真实性标签。在10个主流基线上的大量实验表明,R-T2V达到当前最佳性能,准确率超越第二好基线12.20个百分点,宏F1提升8.46个百分点。
原文摘要 · Abstract (English)
Recent text-to-video (T2V) generation models enable fake news videos to be synthesized from scratch, shifting the threat beyond cheap fakes assembled from existing footage. Such news videos can closely match fabricated narratives, creating a modality alignment trap for existing detectors. Existing datasets lack pure synthesis fake news videos. Although directly prompting T2V models with descriptions of fake news videos can yield perfectly aligned samples, it reduces the fake news video detection (FNVD) to unimodal shortcuts and causes semantic-visual degeneration. To counter this, we formulate T2V-FNVD as a novel ternary classification task with three labels (real, cheap fake, and pure synthesis fake) and construct the first pure synthesis fake news video dataset (PS-FNVD). PS-FNVD includes fabricated events with aligned deception (Type 1) and true events with false visual provenance (Type 2), preventing models from exploiting unimodal shortcuts. Furthermore, we propose the Reasoning-guided T2V-FNVD (R-T2V) framework. Trained through conditioned rationale generation and supervised fine-tuning, R-T2V integrates high-level semantic logic with low-level physical generative traces to predict the ternary veracity label. Extensive experiments across 10 prevailing baselines show that R-T2V achieves the state-of-the-art performance, outperforming the second-best baseline by 12.20 percentage points in accuracy and 8.46 percentage points in macro $F_1$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。