面向人脸编辑的跨模态语义通信,大幅降低传输带宽。
Editable-DeepSC: Reliable Cross-Modal Semantic Communications for Facial Editing
- 将编辑任务嵌入通信链,通过联合编码提升语义信息保留。
- 在高分辨率和分布外场景下仍保持高质量编辑效果,节省传输带宽。
- 适合社交媒体中实时人脸编辑等交互式视觉应用。
交互式计算机视觉在现实应用中至关重要,其性能高度依赖通信网络。然而,传统通信的数据导向特性难以满足交互式视觉任务需求。近年来兴起的语义通信仅传输与任务相关的语义信息,展现出良好前景。但针对社交媒体中重要的语义人脸编辑任务,相关通信挑战仍未充分探索。本文提出可编辑的深度语义通信(Editable-DeepSC),一种新型跨模态语义通信方法。首先理论分析了分离通信与编辑的方案,强调通过迭代属性匹配实现联合编辑-信道编码(JECC)的必要性,以将编辑融入通信链,更好地保留语义互信息。为紧凑表示高维数据,利用预训练的StyleGAN先验进行反演编码。针对动态信道噪声,提出基于信噪比感知的模型微调通道编码。大量实验表明,Editable-DeepSC在高分辨率和分布外(OOD)设置下仍能实现优异编辑效果,并显著减少传输带宽。
原文摘要 · Abstract (English)
Interactive computer vision (CV) plays a crucial role in various real-world applications, whose performance is highly dependent on communication networks. Nonetheless, the data-oriented characteristics of conventional communications often do not align with the special needs of interactive CV tasks. To alleviate this issue, the recently emerged semantic communications only transmit task-related semantic information and exhibit a promising landscape to address this problem. However, the communication challenges associated with Semantic Facial Editing, one of the most important interactive CV applications on social media, still remain largely unexplored. In this paper, we fill this gap by proposing Editable-DeepSC, a novel cross-modal semantic communication approach for facial editing. Firstly, we theoretically discuss different transmission schemes that separately handle communications and editings, and emphasize the necessity of Joint Editing-Channel Coding (JECC) via iterative attributes matching, which integrates editings into the communication chain to preserve more semantic mutual information. To compactly represent the high-dimensional data, we leverage inversion methods via pre-trained StyleGAN priors for semantic coding. To tackle the dynamic channel noise conditions, we propose SNR-aware channel coding via model fine-tuning. Extensive experiments indicate that Editable-DeepSC can achieve superior editings while significantly saving the transmission bandwidth, even under high-resolution and out-of-distribution (OOD) settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。