为提升模型准确率设计的多摄像头视频压缩框架
DMVC: Multi-Camera Video Compression Network aimed at Improving Deep Learning Accuracy
- 按批量处理多路视频,保留对模型关键的语义信息
- 在显著压缩数据的同时保持甚至提升模型精度
- 适合智能城市、自动驾驶等需高效视频分析的场景
我们提出一种面向深度学习应用的新型视频压缩框架DMVC,突破传统以人眼感知为主的压缩思路,专注于保留对机器学习任务至关重要的语义信息,同时大幅降低数据量。该框架支持批量处理多路视频流,具备良好的可扩展性与处理效率,并提供轻量级(实时响应)与高精度(高准确率)两种重建模式。基于设计的深度学习算法,能有效分离关键信息与冗余内容,确保输入模型的数据具有最高相关性。实验基于城市监控与自动驾驶等多种数据集验证,结果表明,DMVC在实现显著数据压缩的同时,仍能维持甚至提升下游机器学习任务的准确率。该工作为视频压缩与深度学习融合树立了新标杆,有望推动智能城市与自主系统等领域的发展。
原文摘要 · Abstract (English)
We introduce a cutting-edge video compression framework tailored for the age of ubiquitous video data, uniquely designed to serve machine learning applications. Unlike traditional compression methods that prioritize human visual perception, our innovative approach focuses on preserving semantic information critical for deep learning accuracy, while efficiently reducing data size. The framework operates on a batch basis, capable of handling multiple video streams simultaneously, thereby enhancing scalability and processing efficiency. It features a dual reconstruction mode: lightweight for real-time applications requiring swift responses, and high-precision for scenarios where accuracy is crucial. Based on a designed deep learning algorithms, it adeptly segregates essential information from redundancy, ensuring machine learning tasks are fed with data of the highest relevance. Our experimental results, derived from diverse datasets including urban surveillance and autonomous vehicle navigation, showcase DMVC's superiority in maintaining or improving machine learning task accuracy, while achieving significant data compression. This breakthrough paves the way for smarter, scalable video analysis systems, promising immense potential across various applications from smart city infrastructure to autonomous systems, establishing a new benchmark for integrating video compression with machine learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。