用语义通信提升移动端3D重建实时性与鲁棒性
Toward Semantic Communication for Real-time Mobile 3D Reconstruction

- 传输时附带像素级置信图,量化区域可靠性
- 置信度引导下姿态估计误差降低42%,3D结构更一致
- 适合无线信道不稳定场景下的实时三维重建
实时移动端3D重建对自动驾驶和数字孪生等应用至关重要,移动平台持续采集图像流并传至服务器进行场景理解。与离线重建不同,相机位姿与场景几何需在采集过程中实时估计,多视角一致性成为关键要求,且几何估计对通信畸变高度敏感。语义通信(SemCom)通过传输紧凑的语义信息,在不可靠链路下保留任务关键数据,具有潜力。然而现有方法多针对单图像或视图级优化,缺乏对几何估计的显式可靠性支持,限制了其在实时移动端3D重建中的应用。为此,本文提出一种面向实时移动端3D重建的语义通信框架。该框架包含一个语义收发器,输出重建图像及像素级置信图,量化各区域可靠性。进一步引入置信度引导的几何估计方法,将置信度融入基于RANSAC的姿态初始化与捆绑调整,减少不可靠区域影响,提升噪声信道下的鲁棒性。仿真结果表明,相比现有语义通信和传统编解码分离方案,本框架在保持高图像质量的同时,显著提升姿态估计精度与3D结构一致性。
原文摘要 · Abstract (English)
Real-time mobile 3D reconstruction is fundamental to many emerging applications such as autonomous navigation and digital twin construction, where a moving platform continuously captures an image stream and transmit to a computing server for scene understanding. Unlike offline reconstruction, camera poses and scene geometry are estimated on-the-fly during acquisition, making multi-view consistency a real-time requirement and rendering geometric estimation highly sensitive to communication-induced distortions. Semantic communication (SemCom) transmits compact semantic information, offering a promising way to preserve task-critical data over unreliable links. However, existing designs are optimized at the image or single-view level and without providing explicit reliability information for geometric estimation, limiting their applicability to real-time mobile 3D reconstruction. In this context, we propose a SemCom framework for real-time mobile 3D reconstruction. The framework includes a semantic transceiver that outputs a reconstructed image alongside a pixel-wise confidence map, quantifying the reliability of each region. We further introduce a confidence-guided geometric estimation method, incorporating confidence into RANSAC-based pose initialization and bundle adjustment to reduce the influence of unreliable regions and enhance robustness under noisy channels. Simulations show that, compared to existing SemCom and traditional seperate source and channel coding, our framework maintains high image quality while significantly improving pose estimation accuracy and 3D structural consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。