通过不确定性感知提升多模态边缘推理通信效率
Communication-Efficient Multi-Modal Edge Inference via Uncertainty-Aware Distributed Learning
- 分三阶段分布式学习,本地自监督编码减少通信开销
- 融合模态置信度,应对无线信道噪声与模态退化
- 仅对不确定样本请求额外特征,优化通信与精度平衡
语义通信正成为分布式边缘智能的关键使能技术,因其能够传递任务相关语义。然而,在带宽受限的无线链路上实现高效训练与鲁棒推理仍具挑战,尤其在多模态边缘推理(MMEI)中面临双重难题:1)因系统具多模态特性,分布式学习带来高昂通信开销;2)在信道变化与多模态输入噪声下鲁棒性不足。本文提出一种三阶段通信感知的分布式学习框架,提升训练与推理效率并保持无线信道上的鲁棒性。第一阶段,设备执行本地多模态自监督学习,构建共享与模态特定编码器,无需设备-服务器通信,降低通信成本;第二阶段,采用中心化证据融合的分布式微调,校准各模态不确定性,可靠聚合受噪声或信道衰落影响的特征;第三阶段,基于不确定性的反馈机制,仅对不确定样本请求额外特征,优化分布式场景下的通信-精度权衡。在RGB-深度室内场景分类实验中,该框架以更少的训练通信轮次达到更高准确率,并在模态退化或信道变化下保持鲁棒性,优于现有自监督与全监督基线方法。
原文摘要 · Abstract (English)
Semantic communication is emerging as a key enabler for distributed edge intelligence due to its capability to convey task-relevant meaning. However, achieving communication-efficient training and robust inference over wireless links remains challenging. This challenge is further exacerbated for multi-modal edge inference (MMEI) by two factors: 1) prohibitive communication overhead for distributed learning over bandwidth-limited wireless links, due to the \emph{multi-modal} nature of the system; and 2) limited robustness under varying channels and noisy multi-modal inputs. In this paper, we propose a three-stage communication-aware distributed learning framework to improve training and inference efficiency while maintaining robustness over wireless channels. In Stage~I, devices perform local multi-modal self-supervised learning to obtain shared and modality-specific encoders without device--server exchange, thereby reducing the communication cost. In Stage~II, distributed fine-tuning with centralized evidential fusion calibrates per-modality uncertainty and reliably aggregates features distorted by noise or channel fading. In Stage~III, an uncertainty-guided feedback mechanism selectively requests additional features for uncertain samples, optimizing the communication--accuracy tradeoff in the distributed setting. Experiments on RGB--depth indoor scene classification show that the proposed framework attains higher accuracy with far fewer training communication rounds and remains robust to modality degradation or channel variation, outperforming existing self-supervised and fully supervised baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。