arXiv:2604.23247cs.CV2026-04中稿 · TrustFA Workshop, …

无需预处理,通过帧间特征差分实现驱动者身份指纹识别

Micro-Expression-Aware Avatar Fingerprinting via Inter-Frame Feature Differencing

论文配图:Micro-Expression-Aware Avatar Fingerprinting via Inter-Frame Feature Differencing
图 1 · 摘自论文原文
  • 直接在原始视频上使用微表情感知的骨干网络,通过帧间特征相减提取运动动态
  • 在NVFAIR数据集上达到0.877的AUC,多数生成器对下表现优于基于关键点的方法
  • 首次证明运动动态是区分驱动者的核心,外观特征反而会干扰识别

虚拟头像指纹识别——验证合成语音头像的驱动者身份而非真实性——是保障人脸重演技术合规使用的关键。现有方法依赖固定且不可微的关键点提取阶段,无法实现从原始像素端到端优化。本文提出一种无需预处理的系统,采用微表情感知的骨干网络,以帧间特征差分为核心设计:连续特征图在学习的深层特征空间中相减,使时间上稳定的外观成分输出为零,而驱动者特有的运动动态得以保留。在NVFAIR上的受控消融实验表明,时间运动贡献了绝大部分判别性能,而原始外观特征反而会降低身份区分能力。骨干网络选择与差分原则均至关重要:仅用通用编码器时,外观主导的特征在相邻帧间趋同,差分效果失效;而微表情感知的F5C骨干能保持可辨识的运动变化,使差分操作有效。本模型不依赖外部预处理,在NVFAIR上整体AUC达0.877,多数跨生成器对下性能匹配或超越基于关键点的基线。

原文摘要 · Abstract (English)

Avatar fingerprinting, i.e., verifying who drives a synthetic talking-head video rather than whether it is real, is a critical safeguard for authorized use of face-reenactment technology. Existing methods rely on a fixed, non-differentiable landmark extraction stage that prevents the fingerprinting model from being optimized end-to-end from raw pixels. We propose a preprocessing-free system built on a micro-expression-aware backbone operating on raw video frames, with inter-frame feature differencing as the core design principle: consecutive feature maps are subtracted in the learned deep feature space, so that temporally stable appearance dimensions contribute zero to the output while driver-specific motion dynamics are preserved. A controlled ablation on NVFAIR confirms that temporal motion accounts for the large majority of discriminative performance, and that raw appearance features actively degrade identity separation. Both the choice of backbone and the differencing principle are essential: differencing alone is insufficient when applied to a generic encoder, as appearance-dominated features collapse to near-identical representations across adjacent frames, while the micro-expression-aware F5C backbone retains measurable motion variation that the differencing operation can exploit. Without any external preprocessing, our model achieves an overall AUC of 0.877 on NVFAIR and matches or exceeds the landmark-based baseline on the majority of cross-generator pairs.

身份识别生成内容检测视频指纹微表情

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。