用视觉语言模型指导物理模型,统一去除云层干扰
Physics-Guided VLM Priors for All-Cloud Removal
- 将VLM语义先验转为散射参数和可信度图,实现自适应修复
- 在透光区用物理模型保辐射精度,在遮挡区用时序参考重建
- 无需手动分云类型,适合复杂混合云场景
云去除是光学遥感中的基础挑战,因云层异质性导致退化。薄云通过部分透射扭曲辐射,厚云则遮挡地表。现有方法将薄云校正与厚云重建分离,需显式判断云类型,常导致误差累积和混合云场景中的不连续。为此提出物理引导的视觉语言模型统一去云方法(PhyVLM-CR),将视觉语言模型(如Qwen)的认知先验转化为物理散射参数与幻觉可信度图。利用该可信度图作为连续软门控,自适应加权:在高透射区域优先采用物理反演以保持辐射保真度,低置信度遮挡区域无缝切换至时序参考重建。该机制无需显式边界分割,确保异质云覆盖下的连贯去云效果。在真实Sentinel-2地表反射率影像上的实验表明,本方法在去云与内容保持间取得显著平衡,实现无幻觉结果,定量精度大幅优于现有方法。
原文摘要 · Abstract (English)
Cloud removal is a fundamental challenge in optical remote sensing due to the heterogeneous degradation. Thin clouds distort radiometry via partial transmission, while thick clouds occlude the surface. Existing pipelines separate thin-cloud correction from thick-cloud reconstruction, requiring explicit cloud-type decisions and often leading to error accumulation and discontinuities in mixed-cloud scenes. Therefore, a novel approach named Physical-VLM All-Cloud Removal (PhyVLM-CR) that integrates the semantic capability of Vision-Language Model (VLM) into a physical restoration model, achieving high-fidelity unified cloud removal. Specifically, the cognitive prior from a VLM (e.g., Qwen) is transformed into physical scattering parameters and a hallucination confidence map. Leveraging this confidence map as a continuous soft gate, our method achieves a unified restoration via adaptive weighting: it prioritizes physical inversion in high-transmission regions to preserve radiometric fidelity, while seamlessly transitioning to temporal reference reconstruction in low-confidence occluded areas. This mechanism eliminates the need for explicit boundary delineation, ensuring a coherent removal across heterogeneous cloud covers. Experiments on real-world Sentinel-2 surface reflectance imagery confirm that our approach achieves a remarkable balance between cloud removal and content preservation, delivering hallucination-free results with substantially improved quantitative accuracy compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。