利用手机自拍时的微动信号,可有效识别深度伪造和视频注入攻击。
Selfie-Capture Dynamics as an Auxiliary Signal Against Deepfakes and Injection Attacks for Mobile Identity Verification

- 通过分析自拍时的传感器运动轨迹,构建辅助验证信号。
- 最低误接受率32.0%,在2.37%误拒绝率下全拒静止攻击。
- 适合用于移动端身份认证中的防伪与用户识别,尤其关注低摩擦体验。
移动远程身份验证(RIdV)系统易受面部视频流篡改攻击,包括演示攻击、实时深度伪造和视频注入。欧洲标准如ETSI TS 119 461和CEN/TS 18099要求在摄像头基础上引入互补证据通道。本文研究自拍过程中记录的被动运动痕迹是否可作为伪造检测与用户验证的辅助信号。提出CanSelfie数据集,包含30名参与者在商用RIdV应用中以50Hz采样采集的375组多传感器真实序列,以及静态、手持和时间偏移的攻击模拟场景。在不同传感器配置与时间窗口下,对比7种多变量时间序列分类器和8种整体序列异常检测方法。对于伪造检测,仅使用加速度计的ROCKAD实现0.00%误拒绝率和43.8%误接受率;QUANT+3-NN在2.37%误拒绝率下达到最低整体误接受率32.0%,并成功拒绝对所有静止攻击代理。对于同设备同会话用户验证,WEASEL+MUSE使用9个传感器通道达到1.07%等错误率(EER)。分析表明,原始加速度数据因保留重力与方向信息最具信息量;封闭集分类准确率不能代表验证性能,因阈值校准依赖于得分分布。结果表明,短时自拍运动痕迹蕴含可测量的伪造与身份相关特征,支持其作为低摩擦辅助信号,同时指出需开展跨设备、跨会话及真实注入攻击评估。
原文摘要 · Abstract (English)
Mobile remote identity verification (RIdV) systems are exposed to attacks that manipulate or replace the facial video stream, including presentation attacks, real-time deepfakes, and video injection. Recent European requirements, including ETSI TS 119 461 and CEN/TS 18099, motivate complementary evidence channels beyond camera-based presentation-attack detection. This paper investigates whether passive motion traces recorded during selfie capture provide auxiliary evidence for spoof screening and user verification. We introduce CanSelfie, a dataset of 375 bona fide multi-sensor sequences collected at 50\,Hz from 30 participants using a commercial mobile RIdV application, together with stationary, handheld, and temporally shifted attack-proxy scenarios. We benchmark 7 multivariate time-series classifiers and 8 whole-series anomaly detectors across sensor configurations and temporal windows. For spoof screening, accelerometer-only ROCKAD obtains 0.00\% false rejection rate (FRR) and 43.8\% false acceptance rate (FAR), while QUANT+3-NN obtains the lowest overall FAR of 32.0\% at 2.37\% FRR; both reject all stationary attack proxies. For same-device and same-session user verification, WEASEL+MUSE reaches 1.07\% equal error rate (EER) using 9 sensor channels. The analysis shows that raw accelerometer data, preserving gravity and orientation cues, is the most informative modality, and that closed-set classification accuracy alone does not imply good verification performance because threshold calibration depends on score distributions. The findings suggest that short selfie-capture motion traces contain measurable spoof-related and identity-related information, supporting their use as a low-friction auxiliary signal while also identifying the need for cross-device, cross-session, and real injection-attack evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。