通过消除模态冗余提升多模态通信的可靠性与效率
Robust Multi-modal Task-oriented Communications with Redundancy-aware Representations
- 分两阶段压缩各模态并抑制跨模态冗余
- 在低信噪比下仍保持高情感识别准确率
- 适合语音视频融合任务中的可靠传输场景
面向多模态数据的语义通信可在噪声大、带宽受限的信道中高效传输任务相关信息。然而,如何同时压缩模态间冗余并提升语义可靠性仍是关键挑战。为此,我们提出一种鲁棒高效的多模态任务导向通信框架,结合两阶段变分信息瓶颈(VIB)与互信息(MI)冗余最小化机制。第一阶段对文本、音频、视频等单模态分别应用VIB,保留任务特异性特征;随后引入对抗训练的MI最小化模块,抑制跨模态依赖,促进互补性而非冗余。第二阶段使用多模态VIB融合表示,增强抗信道失真能力。在多模态情感识别任务上的实验表明,该框架在低信噪比条件下显著优于现有基线,在准确率与可靠性上均有提升。本工作提供了一个联合优化模态特定压缩、模态间冗余与通信可靠性的原则性框架。
原文摘要 · Abstract (English)
Semantic communications for multi-modal data can transmit task-relevant information efficiently over noisy and bandwidth-limited channels. However, a key challenge is to simultaneously compress inter-modal redundancy and improve semantic reliability under channel distortion. To address the challenge, we propose a robust and efficient multi-modal task-oriented communication framework that integrates a two-stage variational information bottleneck (VIB) with mutual information (MI) redundancy minimization. In the first stage, we apply uni-modal VIB to compress each modality separately, i.e., text, audio, and video, while preserving task-specific features. To enhance efficiency, an MI minimization module with adversarial training is then used to suppress cross-modal dependencies and to promote complementarity rather than redundancy. In the second stage, a multi-modal VIB is further used to compress the fused representation and to enhance robustness against channel distortion. Experimental results on multi-modal emotion recognition tasks demonstrate that the proposed framework significantly outperforms existing baselines in accuracy and reliability, particularly under low signal-to-noise ratio regimes. Our work provides a principled framework that jointly optimizes modality-specific compression, inter-modal redundancy, and communication reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。