用物理模型约束视觉变压器,提升遥感图像重建效果
Knowledge-Guided Masked Autoencoder with Linear Spectral Mixing and Spectral-Angle-Aware Reconstruction
- 将线性光谱混合模型和光谱角映射作为物理约束嵌入自编码器
- 在有限监督下实现更高重建精度与训练稳定性
- 适合遥感、地质等需物理可解释性的领域应用
将领域知识融入深度学习已成为提升模型可解释性、泛化能力与数据效率的有前景方向。本文提出一种基于视觉变压器的知识引导掩码自编码器,在自监督重建过程中引入科学先验知识。该方法不依赖纯数据驱动优化,而是将线性光谱混合模型(LSMM)作为物理约束,并结合基于物理的光谱角映射(SAM),确保学习到的表征符合观测信号与其潜在组分之间的已知结构关系。框架联合优化LSMM损失、SAM损失与传统Huber损失,兼顾特征空间中的数值精度与几何一致性。该知识引导设计提升了重建保真度,在弱监督条件下稳定训练,并生成基于物理原理的可解释潜在表征。实验表明,所提模型显著改善重建质量并提升下游任务性能,凸显了在基于Transformer的自监督学习中嵌入物理启发归纳偏置的潜力。
原文摘要 · Abstract (English)
Integrating domain knowledge into deep learning has emerged as a promising direction for improving model interpretability, generalization, and data efficiency. In this work, we present a novel knowledge-guided ViT-based Masked Autoencoder that embeds scientific domain knowledge within the self-supervised reconstruction process. Instead of relying solely on data-driven optimization, our proposed approach incorporates the Linear Spectral Mixing Model (LSMM) as a physical constraint and physically-based Spectral Angle Mapper (SAM), ensuring that learned representations adhere to known structural relationships between observed signals and their latent components. The framework jointly optimizes LSMM and SAM loss with a conventional Huber loss objective, promoting both numerical accuracy and geometric consistency in the feature space. This knowledge-guided design enhances reconstruction fidelity, stabilizes training under limited supervision, and yields interpretable latent representations grounded in physical principles. The experimental findings indicate that the proposed model substantially enhances reconstruction quality and improves downstream task performance, highlighting the promise of embedding physics-informed inductive biases within transformer-based self-supervised learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。