用新模型提升VR点云语义识别,精度93.2%同时大幅降算力和内存
ESP-PCT: Enhanced VR Semantic Performance through Efficient Compression of Temporal and Spatial Redundancies in Point Cloud Transformers
- 分两阶段联合训练,压缩时空冗余提升效率
- 精度达93.2%,计算量降低76.9%,内存减少78.2%
- 适合需要高精度低资源的VR语义识别场景
语义识别对虚拟现实(VR)应用至关重要,可实现沉浸式交互体验。一种有前景的方法是利用毫米波(mmWave)信号生成点云。然而,现有mmWave点云模型存在高计算与内存开销问题,影响其效率与可靠性。为此,本文提出ESP-PCT,一种面向VR应用的增强型语义性能点云变换器,采用两阶段语义识别框架。ESP-PCT结合传感点云数据的高精度特性,优化语义识别流程,其中定位与关注阶段以端到端方式联合训练。我们在多种VR语义识别条件下评估了ESP-PCT,显著提升了识别效率。值得注意的是,相比现有Point Transformer模型,ESP-PCT在保持93.2%精度的同时,将计算需求(FLOPs)降低76.9%,内存使用减少78.2%。这凸显了其在VR语义识别中的潜力,实现了高精度与冗余消除的兼顾。项目代码与数据已公开于https://github.com/lymei-SEU/ESP-PCT。
原文摘要 · Abstract (English)
Semantic recognition is pivotal in virtual reality (VR) applications, enabling immersive and interactive experiences. A promising approach is utilizing millimeter-wave (mmWave) signals to generate point clouds. However, the high computational and memory demands of current mmWave point cloud models hinder their efficiency and reliability. To address this limitation, our paper introduces ESP-PCT, a novel Enhanced Semantic Performance Point Cloud Transformer with a two-stage semantic recognition framework tailored for VR applications. ESP-PCT takes advantage of the accuracy of sensory point cloud data and optimizes the semantic recognition process, where the localization and focus stages are trained jointly in an end-to-end manner. We evaluate ESP-PCT on various VR semantic recognition conditions, demonstrating substantial enhancements in recognition efficiency. Notably, ESP-PCT achieves a remarkable accuracy of 93.2% while reducing the computational requirements (FLOPs) by 76.9% and memory usage by 78.2% compared to the existing Point Transformer model simultaneously. These underscore ESP-PCT's potential in VR semantic recognition by achieving high accuracy and reducing redundancy. The code and data of this project are available at \url{https://github.com/lymei-SEU/ESP-PCT}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。