arXiv:2606.17809cs.CV2026-06被引 1

构建百万级花粉显微数据集,实现跨地域精准识别与结构化描述。

Million-scale multimodal pollen microscopy with expert-guided foundation models

论文配图:Million-scale multimodal pollen microscopy with expert-guided foundation models
图 1 · 摘自论文原文
  • 基于专家引导的多模态模型,自动提取花粉图像特征并生成描述。
  • 检测精度达99.6%,跨区域检索准确率仍保持0.811(对比基线0.262)。
  • 适合花粉分类、跨域适应与显微医学影像研究者使用。

自动化花粉识别在气溶胶生物学、古生态学与生物多样性监测中仍是瓶颈,因可扩展系统需在不同制样方式、扫描设备和地理来源下保持泛化性,同时保留花粉学可解释性。为此,我们构建了百万级多模态花粉显微资源Pollen AI Atlas,涵盖四个地理来源、四种扫描设置、45个属及1个仅含科级分类单元的样本,覆盖31个植物科。以每张原始切片手动选取一个代表性颗粒为种子,通过标记级挖掘与过滤,获得1,511,390个释放的花粉粒检测,专家校验区域中提案精度达99.6%。每个检测均配以五种开源视觉-语言模型生成的粒级形态描述,由专家验证的花粉学锚点引导,涵盖孔口系统、壁面纹饰、形状与大小等结构信息。在所评估模型中,Gemma4生成的主描述集控制最佳,兼具长度稳定、无分类名或尺寸泄露,且文本检索性能最强。基于冻结视觉特征的基准测试达到88.16%的top-1准确率;跨区域检索显示,当图像相似性下降时,基于描述的文本嵌入仍保持稳健(mAP@20为0.811对比0.262)。公开数据、标注、描述、划分、代码与权重为花粉识别、跨区域域适应及领域特定多模态显微学习提供基准。

原文摘要 · Abstract (English)

Automated pollen identification from microscopy remains a bottleneck in aerobiology, palaeoecology and biodiversity monitoring, because scalable systems must generalise across specimen preparation, scanner settings and geographic origins while retaining palynological interpretability. To address this gap, we present a million-scale multimodal pollen microscopy resource, Pollen AI Atlas, assembled from pure-species whole-slide bright-field images spanning four geographic origins, four scanner settings, 45 genera, and one family-only taxon across 31 botanical families. Seeded by one manually selected exemplar per source slide, token-level mining and filtering produced 1,511,390 released grain detections with 99.6\% proposal precision in expert-curated test regions. Each detection was paired with machine-generated grain-level morphological captions from five open-weight vision--language models, guided by expert-verified palynological anchors, yielding structured descriptions of aperture systems, wall ornamentation, shape and size. Among the evaluated models, Gemma4 provided the most controlled primary caption set, combining tight length control, no detected taxon-name or numeric-size leakage and the strongest text-retrieval performance. Baseline benchmarks with frozen visual features reached 88.16\% top-1 accuracy, while cross-regional retrieval showed that caption-derived text embeddings remained robust when image similarity degraded (mAP@20 0.811 versus 0.262). Released data, annotations, captions, splits, code, and weights provide a benchmark for pollen recognition, cross-regional domain adaptation and domain-specific multimodal microscopy learning.

花粉识别多模态显微图像领域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。