arXiv:2510.14241cs.CV2025-10ICCV被引 9

通过音素-时序与身份动态分析,精准识别高阶伪造视频中的微小异常。

PIA: Deepfake Detection Using Phoneme-Temporal and Identity-Dynamic Analysis

论文配图:PIA: Deepfake Detection Using Phoneme-Temporal and Identity-Dynamic Analysis
图 1 · 摘自论文原文
  • 融合音素、唇部几何与人脸身份嵌入,实现多模态联合检测
  • 在Deepfake Detection Challenge数据集上达到93.2%准确率,优于现有方法
  • 适合需要高精度伪造视频检测的安防与内容审核场景

伪造媒体的兴起使深度伪造成为严重威胁,涉及唇形同步修改、换脸及驱动型面部合成等多种生成技术。传统检测方法依赖人工设计的音素-视觉对应阈值、基础帧级一致性检查或单一模态策略,难以识别由GAN、扩散模型和神经渲染等先进生成模型产生的现代深度伪造。这些技术虽能生成近乎完美的单帧画面,却常产生人眼难察的微小时间差异。本文提出一种新型多模态音视频框架——音素-时序与身份动态分析(PIA),融合语言、动态面部运动与面部识别线索,利用音素序列、唇部几何数据与先进的人脸身份嵌入,通过多模态互补信息显著提升对细微伪造痕迹的检测能力。代码已开源。

原文摘要 · Abstract (English)

The rise of manipulated media has made deepfakes a particularly insidious threat, involving various generative manipulations such as lip-sync modifications, face-swaps, and avatar-driven facial synthesis. Conventional detection methods, which predominantly depend on manually designed phoneme-viseme alignment thresholds, fundamental frame-level consistency checks, or a unimodal detection strategy, inadequately identify modern-day deepfakes generated by advanced generative models such as GANs, diffusion models, and neural rendering techniques. These advanced techniques generate nearly perfect individual frames yet inadvertently create minor temporal discrepancies frequently overlooked by traditional detectors. We present a novel multimodal audio-visual framework, Phoneme-Temporal and Identity-Dynamic Analysis(PIA), incorporating language, dynamic face motion, and facial identification cues to address these limitations. We utilize phoneme sequences, lip geometry data, and advanced facial identity embeddings. This integrated method significantly improves the detection of subtle deepfake alterations by identifying inconsistencies across multiple complementary modalities. Code is available at https://github.com/skrantidatta/PIA

深度伪造检测多模态分析音视频一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。