通过边缘-云协同重建,保留遥感图像关键结构,提升目标检测精度。
Edge-Cloud Collaborative Reconstruction via Structure-Aware Latent Diffusion for Downstream Remote Sensing Perception

- 边缘端分离低频数据与轻量结构先验,降低传输带宽需求。
- 云端利用结构先验引导扩散模型,显著减少生成幻觉并提升感知质量。
- 适用于高压缩比遥感数据,特别适合小目标检测场景。
高分辨率遥感数据的指数级增长面临卫星到地面传输的严重瓶颈。有限的下行带宽迫使采用极端高压缩比,导致下游机器感知任务(如目标检测)所需的关键高频结构细节被不可逆破坏。现有超分辨率技术虽尝试恢复这些细节,但基于回归的方法常导致纹理过度平滑,生成式扩散模型则易引入结构幻觉,误导检测系统。为此,本文提出结构感知潜空间扩散(SALD)框架,一种非对称边缘-云协同超分辨系统。在资源受限的边缘端,系统将图像解耦为高度压缩的低频数据包和轻量级软结构先验,传输该解耦表示以最小化带宽消耗。在强大云端侧,引入结构门控大卷积核(SGLK)模块与语义引导引擎(SGE),嵌入扩散主干网络中,利用传输的结构先验门控大卷积核运算,有效捕捉航拍场景中的长程依赖关系,同时主动抑制生成幻觉。在MSCM与UCMerced数据集上的大量实验表明,即便在极端带宽约束下,SALD仍实现更优的感知质量(LPIPS),并显著提升场景分类与小目标检测的下游性能。
原文摘要 · Abstract (English)
The exponential surge in high-resolution remote sensing data faces a severe bottleneck in satellite-to-ground transmission. Limited downlink bandwidth forces the use of extreme high-ratio compression, which irreversibly destroys high-frequency structural details essential for downstream machine perception tasks like object detection. While current super-resolution techniques attempt to recover these details, regression-based methods often yield over-smoothed textures, and generative diffusion models frequently introduce structural hallucinations that mislead detection systems. To address this trade-off, we propose the Structure-Aware Latent Diffusion (SALD) framework, an asymmetric edge-cloud collaborative SR system. At the resource-constrained edge, the system decouples imagery into a highly compressed low-frequency payload and a lightweight soft structural prior. Transmitting this decoupled representation minimizes bandwidth consumption. On the powerful cloud side, we introduce a Structure-Gated Large Kernel (SGLK) module and a Semantic-Guidance Engine (SGE) within the diffusion backbone. These modules leverage the transmitted structural priors to gate large-kernel convolutions, effectively capturing long-range dependencies inherent in aerial scenes while actively suppressing generative hallucinations. Extensive experiments on both the MSCM and UCMerced datasets demonstrate that, even under extreme bandwidth constraints, SALD achieves superior perceptual quality (LPIPS) and significantly enhances downstream performance in both scene classification and small-target detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。