arXiv:2510.21346cs.CVcs.AI2025-10被引 1

融合局部细节与全局结构,提升复杂环境下苹果叶病识别准确率

CT-CLIP: A Multi-modal Fusion Framework for Robust Apple Leaf Disease Recognition in Complex Environments

  • 用CNN提取病斑细节,Vision Transformer捕捉整体结构关系
  • 自适应融合模块动态结合多尺度特征,应对病斑形态多样性
  • 引入图文对齐学习,在少量样本下仍保持高精度,适合农业场景

在复杂的果园环境中,不同苹果叶病的表型异质性显著,病斑形态和分布差异大,传统多尺度特征融合方法仅依赖卷积神经网络(CNN)提取的多层特征,难以充分建模局部与全局特征间的关系。为此,本文提出一种多分支识别框架CT-CLIP,协同使用CNN提取局部病斑细节特征,Vision Transformer捕获全局结构关系,并通过自适应特征融合模块(AFFM)动态融合两者,实现局部与全局信息的最佳耦合,有效应对病斑形态与分布的多样性。此外,为减少复杂背景干扰并提升少样本条件下的识别精度,提出一种多模态图像-文本学习方法,借助预训练的CLIP权重,实现视觉特征与疾病语义描述的深度对齐。实验结果表明,CT-CLIP在公开苹果病害数据集和自建数据集上分别达到97.38%和96.12%的准确率,优于多个基线方法。该模型在复杂环境下的农作病害识别中表现强劲,为农业自动化诊断提供了创新且实用的解决方案。

原文摘要 · Abstract (English)

In complex orchard environments, the phenotypic heterogeneity of different apple leaf diseases, characterized by significant variation among lesions, poses a challenge to traditional multi-scale feature fusion methods. These methods only integrate multi-layer features extracted by convolutional neural networks (CNNs) and fail to adequately account for the relationships between local and global features. Therefore, this study proposes a multi-branch recognition framework named CNN-Transformer-CLIP (CT-CLIP). The framework synergistically employs a CNN to extract local lesion detail features and a Vision Transformer to capture global structural relationships. An Adaptive Feature Fusion Module (AFFM) then dynamically fuses these features, achieving optimal coupling of local and global information and effectively addressing the diversity in lesion morphology and distribution. Additionally, to mitigate interference from complex backgrounds and significantly enhance recognition accuracy under few-shot conditions, this study proposes a multimodal image-text learning approach. By leveraging pre-trained CLIP weights, it achieves deep alignment between visual features and disease semantic descriptions. Experimental results show that CT-CLIP achieves accuracies of 97.38% and 96.12% on a publicly available apple disease and a self-built dataset, outperforming several baseline methods. The proposed CT-CLIP demonstrates strong capabilities in recognizing agricultural diseases, significantly enhances identification accuracy under complex environmental conditions, provides an innovative and practical solution for automated disease recognition in agricultural applications.

病害识别多模态融合少样本学习农业AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。