用语义掩码降低多视角传输开销,提升资源受限设备的感知效率。
Resource-Efficient Multiview Perception: Integrating Semantic Masking with Masked Autoencoders
- 基于预训练分割模型与可调幂函数,优先保留图像关键区域。
- 高掩码率下仍保持接近顶尖方法的检测跟踪精度。
- 相比随机掩码和基线方法,显著减少传输数据量,适合无人机等设备。
多视角系统已成为现代计算机视觉的关键技术,在场景理解与分析方面表现优异。然而,面对带宽限制与计算约束,尤其在无人机等资源受限的摄像节点上,此类系统面临挑战。本文提出一种基于掩码自编码器(MAE)的通信高效分布式多视角检测与跟踪新方法。通过引入语义引导的掩码策略,结合预训练分割模型与可调幂函数,动态聚焦于图像中信息量更高的区域。该方法与MAE结合,在大幅降低通信开销的同时有效保留关键视觉信息。在虚拟与真实多视角数据集上的实验表明,即使在高掩码率下,性能仍可媲美当前最优方法;所提选择性掩码算法优于随机掩码,在掩码率上升时仍保持更高准确率与精确度。此外,相较基线方法,本方案显著减少传输数据量,实现多视角跟踪性能与通信效率的平衡。
原文摘要 · Abstract (English)
Multiview systems have become a key technology in modern computer vision, offering advanced capabilities in scene understanding and analysis. However, these systems face critical challenges in bandwidth limitations and computational constraints, particularly for resource-limited camera nodes like drones. This paper presents a novel approach for communication-efficient distributed multiview detection and tracking using masked autoencoders (MAEs). We introduce a semantic-guided masking strategy that leverages pre-trained segmentation models and a tunable power function to prioritize informative image regions. This approach, combined with an MAE, reduces communication overhead while preserving essential visual information. We evaluate our method on both virtual and real-world multiview datasets, demonstrating comparable performance in terms of detection and tracking performance metrics compared to state-of-the-art techniques, even at high masking ratios. Our selective masking algorithm outperforms random masking, maintaining higher accuracy and precision as the masking ratio increases. Furthermore, our approach achieves a significant reduction in transmission data volume compared to baseline methods, thereby balancing multiview tracking performance with communication efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。