用物理知识增强深度学习,提升雷达图像识别的效率与可解释性
Knowledge-Informed Neural Network for Complex-Valued SAR Image Recognition
- 引入物理先验的三阶段压缩架构,融合电磁散射特性
- 参数量仅0.7M~0.95M,仍达最先进性能,数据稀缺下泛化强
- 适合需要可解释性与小模型部署的雷达图像分析场景
面向复数合成孔径雷达(CV-SAR)图像识别的深度学习模型在数据有限和域偏移场景下面临表征三难困境:泛化性、可解释性与效率难以兼顾。本文提出知识引导神经网络(KINN),基于新型“压缩-聚合-压缩”架构,通过物理引导压缩阶段,利用新型字典处理器自适应嵌入物理先验,使紧凑展开网络高效提取稀疏且物理可信的特征;后续聚合模块增强表征,最后通过紧凑分类头与自蒸馏实现语义压缩,学习最具任务相关性的判别性嵌入。在五组SAR基准测试中,KINN在两种变体(CNN: 0.7M,ViT: 0.95M)上均实现参数高效的最优性能,在数据稀缺和分布外场景下表现优异,兼具强泛化能力与可解释性,为解决表征三难问题提供新路径,推动可信遥感AI发展。
原文摘要 · Abstract (English)
Deep learning models for complex-valued Synthetic Aperture Radar (CV-SAR) image recognition are fundamentally constrained by a representation trilemma under data-limited and domain-shift scenarios: the concurrent, yet conflicting, optimization of generalization, interpretability, and efficiency. Our work is motivated by the premise that the rich electromagnetic scattering features inherent in CV-SAR data hold the key to resolving this trilemma, yet they are insufficiently harnessed by conventional data-driven models. To this end, we introduce the Knowledge-Informed Neural Network (KINN), a lightweight framework built upon a novel "compression-aggregation-compression" architecture. The first stage performs a physics-guided compression, wherein a novel dictionary processor adaptively embeds physical priors, enabling a compact unfolding network to efficiently extract sparse, physically-grounded signatures. A subsequent aggregation module enriches these representations, followed by a final semantic compression stage that utilizes a compact classification head with self-distillation to learn maximally task-relevant and discriminative embeddings. We instantiate KINN in both CNN (0.7M) and Vision Transformer (0.95M) variants. Extensive evaluations on five SAR benchmarks confirm that KINN establishes a state-of-the-art in parameter-efficient recognition, offering exceptional generalization in data-scarce and out-of-distribution scenarios and tangible interpretability, thereby providing an effective solution to the representation trilemma and offering a new path for trustworthy AI in SAR image analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。