提出快速推理的无监督视频特征提取网络,提升稀疏性与聚类准确率。
Fast Deep Predictive Coding Networks for Videos Feature Extraction without Labels
- 基于动态规划与优化框架,实现模型状态与原因的快速推断
- 在CIFAR-10、Super Mario Bros等数据集上,学习速率与稀疏比更优
- 适用于无需标签的视频目标识别,具备可解释性强优势
受大脑启发的深度预测编码网络(DPCN)通过双向信息流建模视频特征,即使无标签也能有效工作。其依赖于视频场景的过完备描述,但长期以来缺乏高效的稀疏化技术以获得判别性强且鲁棒的词典。目前最优方法为FISTA。本文提出一种新型DPCN,实现内部模型变量(状态与原因)的快速推理,显著提升特征聚类的稀疏性与准确性。所提无监督学习过程受自适应动态规划与极大化-极小化框架启发,其收敛性得到严格分析。在CIFAR-10、Super Mario Bros游戏视频及Coil-100数据集上的实验验证了该方法的有效性,相比以往DPCN版本,在学习速率、稀疏比和特征聚类准确率方面均有提升。由于DPCN本身具备良好理论基础与可解释性,此进展为无需标签的视频目标识别提供了通用解决方案。
原文摘要 · Abstract (English)
Brain-inspired deep predictive coding networks (DPCNs) effectively model and capture video features through a bi-directional information flow, even without labels. They are based on an overcomplete description of video scenes, and one of the bottlenecks has been the lack of effective sparsification techniques to find discriminative and robust dictionaries. FISTA has been the best alternative. This paper proposes a DPCN with a fast inference of internal model variables (states and causes) that achieves high sparsity and accuracy of feature clustering. The proposed unsupervised learning procedure, inspired by adaptive dynamic programming with a majorization-minimization framework, and its convergence are rigorously analyzed. Experiments in the data sets CIFAR-10, Super Mario Bros video game, and Coil-100 validate the approach, which outperforms previous versions of DPCNs on learning rate, sparsity ratio, and feature clustering accuracy. Because of DCPN's solid foundation and explainability, this advance opens the door for general applications in object recognition in video without labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。