用逻辑门网络实现超快超小视频拷贝检测
Efficient Logic Gate Networks for Video Copy Detection

- 用可训练逻辑门替代传统神经网络,生成紧凑逻辑表示
- 推理速度超11000样本/秒,特征尺寸缩小数个数量级
- 适合大规模实时视频检测系统,资源消耗极低
视频拷贝检测需在多种视觉失真下保持鲁棒的相似性估计,并支持超大规模部署。尽管深度神经网络表现优异,但其计算开销和特征尺寸限制了在高吞吐系统中的应用。本文提出基于可微逻辑门网络(LGN)的视频拷贝检测框架,用紧凑的逻辑表示替代传统的浮点特征提取器。方法结合极端帧压缩、二值化预处理与可训练的LGN嵌入模型,学习逻辑运算与连接关系。训练后模型可离散化为纯布尔电路,实现极快且内存高效的推理。在多个数据集折线和不同难度级别上系统评估了不同相似性策略、二值化方案与LGN架构。实验表明,基于LGN的模型在准确率与排序性能上达到或优于现有模型,特征尺寸缩小数个数量级,推理速度超过11,000样本/秒。结果表明,逻辑基模型为可扩展、资源高效视频拷贝检测提供了有前景的替代方案。
原文摘要 · Abstract (English)
Video copy detection requires robust similarity estimation under diverse visual distortions while operating at very large scale. Although deep neural networks achieve strong performance, their computational cost and descriptor size limit practical deployment in high-throughput systems. In this work, we propose a video copy detection framework based on differentiable Logic Gate Networks (LGNs), which replace conventional floating-point feature extractors with compact, logic-based representations. Our approach combines aggressive frame miniaturization, binary preprocessing, and a trainable LGN embedding model that learns both logical operations and interconnections. After training, the model can be discretized into a purely Boolean circuit, enabling extremely fast and memory-efficient inference. We systematically evaluate different similarity strategies, binarization schemes, and LGN architectures across multiple dataset folds and difficulty levels. Experimental results demonstrate that LGN-based models achieve competitive or superior accuracy and ranking performance compared to prior models, while producing descriptors several orders of magnitude smaller and delivering inference speeds exceeding 11k samples per second. These findings indicate that logic-based models offer a promising alternative for scalable and resource-efficient video copy detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。