用视觉先验指导雷达图像转照片,减少噪声导致的错图问题
OSCAR: Optical-aware Semantic Control for Aleatoric Refinement in Sar-to-Optical Translation
- 用光学图像教雷达图像理解语义,提升特征对齐
- 结合文字和局部图像提示,精准控制生成细节
- 考虑噪声不确定性,自动调节修复重点,适合遥感图像生成
合成孔径雷达(SAR)具备全天候成像能力,但将SAR图像转换为逼真光学图像仍是根本性难题。现有方法常受SAR固有的斑点噪声和几何畸变影响,导致语义误判、纹理模糊和结构幻觉。为此,提出一种新的SAR到光学(S2O)图像翻译框架,包含三项核心技术:(i)跨模态语义对齐,通过从光学教师模型中提炼鲁棒语义先验,训练一个光学感知的SAR学生编码器;(ii)语义引导生成控制,采用语义接地的ControlNet,融合类别感知文本提示提供全局上下文与分层视觉提示实现局部空间引导;(iii)不确定性感知目标函数,显式建模随机性不确定性,动态调节重建焦点,有效缓解由斑点噪声引发的模糊性伪影。大量实验表明,该方法在感知质量与语义一致性上优于当前最优方法。
原文摘要 · Abstract (English)
Synthetic Aperture Radar (SAR) provides robust all-weather imaging capabilities; however, translating SAR observations into photo-realistic optical images remains a fundamentally ill-posed problem. Current approaches are often hindered by the inherent speckle noise and geometric distortions of SAR data, which frequently result in semantic misinterpretation, ambiguous texture synthesis, and structural hallucinations. To address these limitations, a novel SAR-to-Optical (S2O) translation framework is proposed, integrating three core technical contributions: (i) Cross-Modal Semantic Alignment, which establishes an Optical-Aware SAR Encoder by distilling robust semantic priors from an Optical Teacher into a SAR Student (ii) Semantically-Grounded Generative Guidance, realized by a Semantically-Grounded ControlNet that integrates class-aware text prompts for global context with hierarchical visual prompts for local spatial guidance; and (iii) an Uncertainty-Aware Objective, which explicitly models aleatoric uncertainty to dynamically modulate the reconstruction focus, effectively mitigating artifacts caused by speckle-induced ambiguity. Extensive experiments demonstrate that the proposed method achieves superior perceptual quality and semantic consistency compared to state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。