用伪造痕迹驱动的压缩方法,让视觉语言模型更快更准地识破造假。
ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models
- 从伪造痕迹出发设计令牌压缩,不靠语义保留关键证据。
- 仅保留10%令牌时仍达90%以上算力节省和顶尖检测性能。
- 适合需要高效高精度多媒体鉴伪的科研与工业场景。
多模态大语言模型通过生成文本推理实现可解释的多媒体鉴伪,但处理密集视觉序列计算开销大,尤其在高分辨率图像和视频上。现有视觉令牌剪枝方法多基于语义,保留显著物体而丢弃背景区域,而伪造痕迹如高频异常和时间抖动常藏于其中。为此,我们提出ForensicZip,一种无需训练的框架,将令牌压缩重构为以伪造为导向的问题。ForensicZip将时间令牌演化建模为带松弛虚拟节点的出生-死亡最优传输问题,量化物理不连续性以指示瞬态生成伪影。其伪造评分结合传输基新颖性与高频先验,在大比例压缩下分离伪造证据与语义内容。在深度伪造与AIGC基准测试中,10%令牌保留率下实现2.97倍加速与超过90%的浮点运算量减少,同时保持最先进检测性能。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) enable interpretable multimedia forensics by generating textual rationales for forgery detection. However, processing dense visual sequences incurs high computational costs, particularly for high-resolution images and videos. Visual token pruning is a practical acceleration strategy, yet existing methods are largely semantic-driven, retaining salient objects while discarding background regions where manipulation traces such as high-frequency anomalies and temporal jitters often reside. To address this issue, we introduce ForensicZip, a training-free framework that reformulates token compression from a forgery-driven perspective. ForensicZip models temporal token evolution as a Birth-Death Optimal Transport problem with a slack dummy node, quantifying physical discontinuities indicating transient generative artifacts. The forensic scoring further integrates transport-based novelty with high-frequency priors to separate forensic evidence from semantic content under large-ratio compression. Experiments on deepfake and AIGC benchmarks show that at 10\% token retention, ForensicZip achieves $2.97\times$ speedup and over 90\% FLOPs reduction while maintaining state-of-the-art detection performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。