用自然语言识别新作物,突破传统农业图像识别的物种限制。
CropVLM: A Domain-Adapted Vision-Language Model for Open-Set Crop Analysis

- 通过领域语义对齐训练视觉语言模型,实现农艺术语与图像特征精准匹配。
- 零样本分类准确率达72.51%,检测新物种性能比基线提升10%以上。
- 无需重新训练即可识别未见作物,适合大规模育种和生物多样性研究。
高通量植物表型分析是现代育种的关键,但受限于“表型瓶颈”——人工数据采集耗时且易受观察者偏差影响。传统封闭集计算机视觉系统需大量物种特异性标注,难以适应多样化的育种群体。为此,我们提出CropVLM,一种通过领域特定语义对齐(DSSA)适配农业领域的视觉语言模型(VLM)。该模型在52,987张涵盖37个物种的野外自然场景图像-标题对上训练,有效将农艺术语映射至细粒度视觉特征。我们进一步提出混合开放集定位网络(HOS-Net),结合CropVLM实现仅依赖自然语言描述检测新作物,无需再训练。该方法摆脱了对物种特异性数据的依赖,为高通量表型分析提供了可扩展解决方案,加速遗传改良并推动可持续农业所需的大规模生物多样性研究。模型权重与完整流程已在GitHub公开。在全面评估中,CropVLM零样本分类准确率达72.51%,优于七种CLIP类基线;其检测管道在CVTCropDet基准上达到49.17 AP50,热带水果物种上达50.73 AP50,显著优于次优方法的34.89和48.58。
原文摘要 · Abstract (English)
High-throughput plant phenotyping, the quantitative measurement of observable plant traits, is critical for modern breeding but remains constrained by a "phenotyping bottleneck," where manual data collection is labor-intensive and prone to observer bias. Conventional closed-set computer vision systems fail to address this challenge, as they require extensive species-specific annotation and lack the flexibility to handle diverse breeding populations. To bridge this gap, we present CropVLM, a Vision-Language Model (VLM) adapted for the agricultural domain via Domain-Specific Semantic Alignment (DSSA). Trained on 52,987 manually selected image-caption pairs covering 37 species in natural field conditions, CropVLM effectively maps agronomic terminology to fine-grained visual features. We further introduce the Hybrid Open-Set Localization Network (HOS-Net), an architecture that integrates CropVLM to enable the detection of novel crops solely from natural language descriptions without retraining. By eliminating the reliance on species-specific training data, CropVLM provides a scalable solution for high-throughput phenotyping, accelerating genetic gain and facilitating large-scale biodiversity research essential for sustainable agriculture. The trained model weights and complete pipeline implementation are publicly available at: [https://github.com/boudiafA/CropVLM](https://github.com/boudiafA/CropVLM). In comprehensive evaluations, CropVLM achieves 72.51% zero-shot classification accuracy, outperforming seven CLIP-style baselines. Our detection pipeline demonstrates superior zero-shot generalization to novel species, achieving 49.17 AP50 on our CVTCropDet benchmark and 50.73 AP50 on tropical fruit species, compared to 34.89 and 48.58 for the next-best method, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。