提升视频拷贝检测抗时间攻击能力,效率更高且更省资源。
Counteracting temporal attacks in Video Copy Detection
- 基于帧间差异局部极大值改进帧选择策略
- 效率提升1.4至5.8倍,推理速度超2倍快
- 在更低存储与算力下仍保持高精度
视频拷贝检测(VCD)在版权保护与内容验证中至关重要,用于识别大规模视频库中的重复与近似重复内容。META AI挑战赛提供了评估前沿方法的基准,双层检测法(Dual-level detection)因其结合视频编辑检测与帧场景检测而成为优胜方案。然而,我们发现其视频编辑检测组件在处理完全复制内容时存在明显不足,且对时间维度攻击敏感。为此,提出基于帧间差异局部极大值的改进帧选择策略,在显著降低计算开销的同时增强对时间篡改的鲁棒性。相比标准1 FPS方法,本方法效率提升1.4至5.8倍;相较于双层检测法,在微平均精确率(μAP)相当的前提下,实现更优的抗时间攻击性能。在表示尺寸减少56%、推理时间提速超过2倍的条件下,更适合实际资源受限场景。
原文摘要 · Abstract (English)
Video Copy Detection (VCD) plays a crucial role in copyright protection and content verification by identifying duplicates and near-duplicates in large-scale video databases. The META AI Challenge on video copy detection provided a benchmark for evaluating state-of-the-art methods, with the Dual-level detection approach emerging as a winning solution. This method integrates Video Editing Detection and Frame Scene Detection to handle adversarial transformations and large datasets efficiently. However, our analysis reveals significant limitations in the VED component, particularly in its ability to handle exact copies. Moreover, Dual-level detection shows vulnerability to temporal attacks. To address it, we propose an improved frame selection strategy based on local maxima of interframe differences, which enhances robustness against adversarial temporal modifications while significantly reducing computational overhead. Our method achieves an increase of 1.4 to 5.8 times in efficiency over the standard 1 FPS approach. Compared to Dual-level detection method, our approach maintains comparable micro-average precision ($μ$AP) while also demonstrating improved robustness against temporal attacks. Given 56\% reduced representation size and the inference time of more than 2 times faster, our approach is more suitable to real-world resource restriction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。