用分层文本描述和噪声优化,提升图像语义通信的还原精度。
TISC: A Text-Driven Image Semantic Communication System for Faithful Reconstruction

- 分层提取图像属性:全局场景、背景、物体级位置与物理特征
- 通过相似度评分选择初始噪声,使重建图像更贴近原图语义
- 适合对图像还原质量要求高的语义通信应用
生成式图像语义通信将图像转换为文本描述,并在接收端通过基于扩散模型的生成方法实现文本到图像的重建,因其极低带宽消耗而受到广泛关注。然而,现有方法在图像到文本(I2T)语义提取和文本到图像(T2I)语义重建环节仍面临两大瓶颈:(i) I2T中语义丢失与失真,整体描述常遗漏细粒度对象属性和空间位置信息,导致生成文本偏离原始图像语义;(ii) T2I中语义忠实度不足,即使使用相同语义准确的文本,不同初始噪声设置也会使扩散模型生成语义一致性不同的图像。为此,我们提出TISC,一种面向高保真重建的文本驱动图像语义通信框架。TISC包含两项关键设计:(1) 树状结构属性语义提取(TSASE),将语义提取分解为全局场景、背景及物体级属性描述,涵盖每个检测物体的空间位置、形状/姿态、颜色、材质等物理属性;(2) 初始噪声优化(INO)机制,根据视觉与语义一致性的综合评分,在发送端选择最优初始噪声种子。多数据集实验表明,TSASE显著提升了物体位置恢复与语义描述保真度,INO参数研究验证了所选噪声配置的有效性。
原文摘要 · Abstract (English)
Generative image semantic communication converts an image into a text description and then performs text-to-image reconstruction at the receiver via diffusion-based generative models. This paradigm has attracted broad attention due to its extremely low bandwidth cost. However, existing methods still face two critical bottlenecks across image-to-text (I2T) semantic extraction at the transmitter and text-to-image (T2I) semantic reconstruction at the receiver: (i) semantic loss and distortion in I2T, where holistic image descriptions may omit fine-grained object attributes and spatial-position information, causing the generated text to deviate from the original image semantics; and (ii) insufficient semantic faithfulness in T2I, where even with the same semantically faithful text description, different initial noise settings may lead diffusion-based reconstruction to produce images with different levels of semantic consistency with the original image. These issues jointly limit the semantic faithfulness of image reconstruction. To address them, we propose TISC, a text-driven image semantic communication framework tailored for faithful reconstruction. TISC incorporates two key designs: (1) Tree-Structured Attribute Semantic Extraction (TSASE), which decomposes semantic extraction into global scene, background, and object-level attribute descriptions, covering spatial position, shape/pose, color, material, and other physical attributes for each detected object; and (2) an Initial Noise Optimization (INO) mechanism, which selects an initial noise seed at the transmitter according to a comprehensive similarity score that jointly considers visual and semantic consistency. Experiments on multiple datasets show that TSASE improves object-position recovery and semantic description faithfulness, while the INO parameter study supports the adopted configuration for noise selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。