自动提取植物标本标签信息,构建图文匹配数据集。
Automated pipeline for herbarium label digitization

- 分步处理:检测组件、定位文字、识别手写印刷混合文本、结构化语义信息。
- 在450份标本上实现字符错误率4.05%-4.10%,语义准确率44.0%-44.5%。
- 适合生物多样性研究与多模态人工智能数据构建者使用。
数字化的植物标本馆藏现已包含超过一亿张可自由访问的标本图像,成为生态学与进化生物学基础问题研究的关键资源。然而,标本标签中编码的丰富元数据(采集者身份、地理地点、采集日期、生态观察)仍难以大规模获取,限制了生物多样性信息学发展及多模态AI所需的图文配对数据集构建。本文提出HERBIOME,一个模块化端到端自动化标签数字化流程,集成YOLOv8组件检测、CRAFT Hezar文字定位、微调后的TrOCR识别混合手写与印刷文本,并利用GPT-4o Mini进行语义元数据结构化。TrOCR基于多源数据集训练(包含CREMMA-AN、PictoCatalogs及专有标本数据R'ecolNat),达到4.05%-4.10%的字符错误率。在450份法语标本上的端到端评估显示,最大窗口相似度(MWS)为0.614-0.618,语义元数据准确率(SMA)为0.440-0.445;混合训练策略提升语义保真度,随机采样则最大化表面相似度,分类学字段仍是主要瓶颈。该方法大幅降低人工转录负担,支持构建忠实反映标本个体特征的图文配对数据集,是下一代多模态生物多样性AI系统的基础。
原文摘要 · Abstract (English)
Digitized herbarium collections, now comprising over 100 million freely accessible specimen images, have become a critical resource for addressing fundamental questions in ecology and evolutionary biology. Yet the rich metadata encoded in herbarium labels (collector identities, geographic localities, collection dates, and ecological observations) remains largely inaccessible at scale, constraining both biodiversity informatics and the construction of specimen-specific image-text corpora for multimodal AI. We present HERBIOME, a modular end-to-end pipeline for automated herbarium label digitization, integrating YOLOv8-based component detection, CRAFT Hezar word-level text localization, fine-tuned TrOCR for recognition of mixed handwritten and printed text, and GPT-4o Mini for semantic metadata structuring into standardized fields. TrOCR was trained on a multi-source dataset combining general transcription corpora (CREMMA-AN, PictoCatalogs) with herbarium-specific data (RéColNat), achieving a Character Error Rate of 4.05-4.10%. End-to-end evaluation on 450 French herbarium specimens, using a dual-metric framework of Maximum Window Similarity (MWS: 0.614-0.618) and Semantic Metadata Accuracy (SMA: 0.440-0.445), reveals that hybrid training strategies improve semantic fidelity while random sampling maximizes surface similarity, with taxonomic fields remaining the principal bottleneck. By automating the extraction of structured metadata from complex, heterogeneous labels, HERBIOME reduces transcription burden, enables the construction of paired image-text datasets that faithfully capture specimen individuality, which is a prerequisite for next-generation multimodal biodiversity AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。