arXiv:2510.08512cs.CVcs.RO2025-10被引 4

用语义场景图压缩点云,98%体积减少仍保结构与语义

Have We Scene It All? Scene Graph-Aware Deep Point Cloud Compression

  • 基于语义场景图将点云分块,用条件编码器生成紧凑表示
  • 在SemanticKITTI和nuScenes上实现最高压缩率,最多减98%数据量
  • 适合多机器人系统,支持定位优化与地图融合任务

高效传输3D点云数据对集中式与分布式多智能体机器系统中的先进感知至关重要,尤其在依赖边缘与云端处理的背景下。然而点云数据庞大复杂,在带宽受限和连接间歇时易导致系统性能下降。本文提出一种基于语义场景图的深度压缩框架:将点云分解为语义一致的片段,通过受特征逐元素线性调制(FiLM)条件控制的语义感知编码器,将其编码为紧凑的潜在表示;采用基于折叠的解码器,结合潜在特征与图节点属性,实现结构精确重建。在SemanticKITTI和nuScenes数据集上的实验表明,该框架达到当前最优压缩率,数据量最多减少98%,同时保持结构与语义保真度。此外,该方法支持多机器人位姿图优化与地图合并等下游应用,轨迹精度与地图对齐效果可媲美原始激光雷达扫描。

原文摘要 · Abstract (English)

Efficient transmission of 3D point cloud data is critical for advanced perception in centralized and decentralized multi-agent robotic systems, especially nowadays with the growing reliance on edge and cloud-based processing. However, the large and complex nature of point clouds creates challenges under bandwidth constraints and intermittent connectivity, often degrading system performance. We propose a deep compression framework based on semantic scene graphs. The method decomposes point clouds into semantically coherent patches and encodes them into compact latent representations with semantic-aware encoders conditioned by Feature-wise Linear Modulation (FiLM). A folding-based decoder, guided by latent features and graph node attributes, enables structurally accurate reconstruction. Experiments on the SemanticKITTI and nuScenes datasets show that the framework achieves state-of-the-art compression rates, reducing data size by up to 98% while preserving both structural and semantic fidelity. In addition, it supports downstream applications such as multi-robot pose graph optimization and map merging, achieving trajectory accuracy and map alignment comparable to those obtained with raw LiDAR scans.

点云压缩语义图多机器人感知系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。