用多粒度自监督学习提升植物病害图像识别精度
PSMamba: Progressive Self-supervised Vision Mamba for Plant Disease Recognition
- 分层师生架构,分别捕捉中尺度病变分布与局部纹理细节
- 在三个数据集上均超越现有自监督方法,尤其在细粒度和域迁移场景表现优异
- 适合农业视觉诊断、病害早期识别等实际应用
自监督学习(SSL)已成为无需人工标注的表征学习强大范式。然而,现有方法多关注全局对齐,难以捕捉植物病害图像中特有的分层、多尺度病灶模式。为此,我们提出PSMamba,一种融合视觉马尔可夫模型(Vision Mamba)高效序列建模能力的渐进式自监督框架,采用双学生分层蒸馏策略。不同于传统单教师-学生设计,PSMamba使用共享全局教师,以及两个专用学生:一个处理中尺度视图以捕捉病灶分布与叶脉结构,另一个聚焦局部视图以提取纹理不规则和早期病灶等细粒度特征。该多粒度监督促进上下文与细节表征的联合学习,一致性损失确保跨尺度对齐。在三个基准数据集上的实验表明,PSMamba持续优于当前最优的SSL方法,在域偏移和细粒度场景下均展现更高准确率与鲁棒性。
原文摘要 · Abstract (English)
Self-supervised Learning (SSL) has become a powerful paradigm for representation learning without manual annotations. However, most existing frameworks focus on global alignment and struggle to capture the hierarchical, multi-scale lesion patterns characteristic of plant disease imagery. To address this gap, we propose PSMamba, a progressive self-supervised framework that integrates the efficient sequence modelling of Vision Mamba (VM) with a dual-student hierarchical distillation strategy. Unlike conventional single teacher-student designs, PSMamba employs a shared global teacher and two specialised students: one processes mid-scale views to capture lesion distributions and vein structures, while the other focuses on local views to capture fine-grained cues such as texture irregularities and early-stage lesions. This multi-granular supervision facilitates the joint learning of contextual and detailed representations, with consistency losses ensuring coherent cross-scale alignment. Experiments on three benchmark datasets show that PSMamba consistently outperforms state-of-the-art SSL methods, delivering superior accuracy and robustness in both domain-shifted and fine-grained scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。