arXiv:2410.08740cs.CVcs.AI2024-10被引 8

自动提取植物标本页信息,提升生物多样性数据采集效率

Hespi: A pipeline for automatically detecting information from hebarium specimen sheets

  • 用计算机视觉检测标本页组件与标签字段
  • 结合OCR/HTR和大模型,准确提取印刷、手写等多类型文字
  • 模块化设计支持定制模型,适用于全球标本馆

标本相关的生物多样性数据对生物、环境和保护科学至关重要。亟需提升从标本图像中高效提取数据的能力,摆脱人工转录。我们开发了'Hespi'(HErbarium Specimen sheet PIpeline),利用先进计算机视觉技术,自动提取标本页原始标签上的预编目数据。Hespi集成两种目标检测模型:一种用于识别标本页各组件,另一种用于定位原始标签中的字段。它可将标签分类为印刷、打字、手写或混合类型,并使用光学字符识别(OCR)和手写文本识别(HTR)进行内容提取。文本随后通过权威物种数据库校正,并由多模态大型语言模型(LLM)进一步优化。Hespi能跨国际标本馆准确检测并提取标本页文字,其模块化架构支持用户训练并集成自定义模型。

原文摘要 · Abstract (English)

Specimen-associated biodiversity data are crucial for biological, environmental, and conservation sciences. A rate shift is needed to extract data from specimen images efficiently, moving beyond human-mediated transcription. We developed `Hespi' (HErbarium Specimen sheet PIpeline) using advanced computer vision techniques to extract pre-catalogue data from primary specimen labels on herbarium specimens. Hespi integrates two object detection models: one for detecting the components of the sheet and another for fields on the primary primary specimen label. It classifies labels as printed, typed, handwritten, or mixed and uses Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) for extraction. The text is then corrected against authoritative taxon databases and refined using a multimodal Large Language Model (LLM). Hespi accurately detects and extracts text from specimen sheets across international herbaria, and its modular design allows users to train and integrate custom models.

计算机视觉标本数字化OCR生物多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。