首次提出时序取证框架,追踪图像生成视频时像素的动态演化轨迹。
Flow of Truth: Proactive Temporal Forensics for Image-to-Video Generation

- 将视频生成理解为像素随时间运动,而非帧合成
- 设计可学习的取证模板,实现跨模型时序追踪
- 适合需要检测深度伪造视频的安全部门与平台
图像到视频(I2V)生成技术快速发展,虽能从单张图像生成逼真视频,但也带来新的取证挑战。与静态图像不同,I2V内容随时间演变,传统仅关注2D像素篡改定位的方法已失效。随着帧间推进,嵌入痕迹会漂移变形,难以定位。为此,我们提出首个面向I2V生成的主动时序取证框架——Flow of Truth。核心挑战在于寻找一个能与生成过程同步演化的取证签名,而生成本质是创造性转化,非确定性重建。我们创新性地将视频生成重定义为‘像素在时间上的运动’。基于此,提出可学习的取证模板,结合模板引导的流模块,解耦运动与图像内容,实现鲁棒的时序追踪。实验表明,该框架在商业与开源I2V模型上均表现优异,显著提升时序取证性能。
原文摘要 · Abstract (English)
The rapid rise of image-to-video (I2V) generation enables realistic videos to be created from a single image but also brings new forensic demands. Unlike static images, I2V content evolves over time, requiring forensics to move beyond 2D pixel-level tampering localization toward tracing how pixels flow and transform throughout the video. As frames progress, embedded traces drift and deform, making traditional spatial forensics ineffective. To address this unexplored dimension, we present **Flow of Truth**, the first proactive framework focusing on temporal forensics in I2V generation. A key challenge lies in discovering a forensic signature that can evolve consistently with the generation process, which is inherently a creative transformation rather than a deterministic reconstruction. Despite this intrinsic difficulty, we innovatively redefine video generation as *the motion of pixels through time rather than the synthesis of frames*. Building on this view, we propose a learnable forensic template that follows pixel motion and a template-guided flow module that decouples motion from image content, enabling robust temporal tracing. Experiments show that Flow of Truth generalizes across commercial and open-source I2V models, substantially improving temporal forensics performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。