构建首个植物显微图像多任务评估基准,揭示现有模型理解能力严重不足。
Benchmarking Vision-Language Models for Microscopic Plant Image Understanding

- 构建涵盖5000+图像的PlantMicro基准,覆盖多种宿主与成像模态
- 模型在病原体分类任务中仅达34.93%准确率,远超随机猜测但表现不佳
- 适合关注植物微观图像理解的生物学家与AI研究者使用
显微成像为研究植物细胞及亚细胞水平的生物学和病理学提供了关键视觉证据。然而,现有视觉-语言模型(VLMs)评测基准主要聚焦于宏观植物图像,微观领域仍缺乏系统研究。为此,我们提出PlantMicro,一个面向显微植物图像的综合性评测基准。该基准整合了超过5000张来自不同宿主、生物领域和成像模态的图像,并设计了一系列互补任务以捕捉显微图像理解的不同方面。为支持这些任务,我们构建了超过9000个问答对,系统评估VLMs的能力。实验表明,当前VLMs在细粒度识别和生物学推理方面存在明显短板。例如,GPT-5在病原体分类任务中仅取得34.93%的准确率,仅略高于随机猜测基线。结果凸显了现有VLMs在理解植物显微图像方面的显著差距。PlantMicro为推动VLMs向可靠、全面的显微级植物理解迈进提供了标准化基础。
原文摘要 · Abstract (English)
Microscopic imaging provides essential visual evidence for studying plant biology and pathology at the cellular and subcellular levels. However, existing benchmarks on vision-language models primarily focus on macroscopic plant imagery, while the microscopic domain remains underexplored. To address this gap, we present PlantMicro, a comprehensive benchmark for evaluating vision-language models (VLMs) in microscopic plant imagery. PlantMicro integrates more than 5,000 images collected across diverse hosts, biological domains, and imaging modalities. Building on this diversity, we design a set of complementary tasks that capture different facets of microscopic image understanding. To support these tasks, we construct over 9,000 VQA pairs that systematically evaluate the capabilities of VLMs. Experiments on PlantMicro show that current VLMs struggle with fine-grained recognition and biologically grounded reasoning. For example, GPT-5 achieves 34.93% accuracy on the pathogen classification task, which is only modestly above the random-guessing baseline. The results highlight a significant gap in current VLMs' ability to comprehend plant microscopic images. PlantMicro provides a standardized foundation for advancing VLMs toward reliable and comprehensive microscopy-level plant understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。