arXiv:2506.03211cs.CVcs.NI2025-06被引 3

用图像辅助点云传输,实现高效实时的点云语义通信。

Channel-adaptive Cross-modal Generative Semantic Communication for Point Cloud Transmission

  • 融合图像与点云的跨模态语义编码,提升压缩效率。
  • 在低信噪比和带宽受限下仍保持高质量重建。
  • 支持模拟传输,适合自动驾驶等实时场景应用。

随着自动驾驶和扩展现实的快速发展,点云(PCs)的高效传输日益重要。本文提出一种新型通道自适应跨模态生成语义通信框架GenSeC-PC,通过融合图像与点云的语义编码器实现高效率压缩,其中图像作为未传输的辅助信息。解码器基于PointDif架构,结合修正的去噪扩散隐式模型,使解码速度达到毫秒级,实现实时通信。为增强鲁棒性并降低系统复杂度,设计了轻量级、非对称的通道自适应联合语义-信道编码结构,仅需编码器获取平均信噪比(SNR)和可用带宽反馈。与现有方法不同,GenSeC-PC利用生成先验,在源点云噪声或缺失情况下仍可可靠重建。更重要的是,它支持完全模拟传输,避免了传统语义通信中对无误辅助信息传输的需求,显著提升压缩效率。仿真结果验证了跨模态语义提取与双指标引导微调的有效性,表明该框架在低信噪比、带宽限制、2D图像数量变化及未见物体等多样条件下均具强鲁棒性。

原文摘要 · Abstract (English)

With the rapid development of autonomous driving and extended reality, efficient transmission of point clouds (PCs) has become increasingly important. In this context, we propose a novel channel-adaptive cross-modal generative semantic communication (SemCom) for PC transmission, called GenSeC-PC. GenSeC-PC employs a semantic encoder that fuses images and point clouds, where images serve as non-transmitted side information. Meanwhile, the decoder is built upon the backbone of PointDif. Such a cross-modal design not only ensures high compression efficiency but also delivers superior reconstruction performance compared to PointDif. Moreover, to ensure robust transmission and reduce system complexity, we design a streamlined and asymmetric channel-adaptive joint semantic-channel coding architecture, where only the encoder needs the feedback of average signal-to-noise ratio (SNR) and available bandwidth. In addition, rectified denoising diffusion implicit models is employed to accelerate the decoding process to the millisecond level, enabling real-time PC communication. Unlike existing methods, GenSeC-PC leverages generative priors to ensure reliable reconstruction even from noisy or incomplete source PCs. More importantly, it supports fully analog transmission, improving compression efficiency by eliminating the need for error-free side information transmission common in prior SemCom approaches. Simulation results confirm the effectiveness of cross-modal semantic extraction and dual-metric guided fine-tuning, highlighting the framework's robustness across diverse conditions, including low SNR, bandwidth limitations, varying numbers of 2D images, and previously unseen objects.

点云传输语义通信跨模态扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。