arXiv:2608.14587cs.AI2026-08

用规则和大模型构建植物性状提取框架,实现可解释的大规模文献数据挖掘。

An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts: A Plant Science Use Case

  • 结合规则解析与多模型协同,分步提取植物性状
  • 从4961个物种中提取55,737条性状,平均每物种9.1条
  • 大模型提升75%性状覆盖度,适合生物信息学研究者

背景:信息检索(IR)近年来融合密集与稀疏表示、大语言模型(LLMs)及专用检索模型,显著提升排序准确率、相关性与跨语言性能。文档布局分析、段落索引与语义知识表示等补充技术通过捕捉细粒度上下文与结构信息进一步增强检索效果。新兴的代理式大模型框架通过规划、迭代推理、工具调用与多智能体协作拓展了应用边界,同时强调严格评估、伦理考量与可信性,确保真实场景中的负责任部署。本文提出一种模块化、基于代理的植物性状提取流程:光学字符识别(OCR)将PDF转为机器可读文本,分割与索引按属种组织内容;规则解析器提取结构化植物性状,大语言模型(LLMs)集成则扩展性状词汇并解决歧义。该方法确保精准物种识别、可扩展标注与可解释的文本描述整合,支持在大规模植物文献中稳健、可解释的数据提取。结果:在三个区域植物数据库上,系统共提取55,737条性状注释,覆盖4,961个物种,平均每物种9.1条性状。大模型增强使75%性状的覆盖度提升,总注释量增加59%。尽管OCR引擎选择对物种识别影响较小,整体注释数量保持稳定,证明该流水线在大规模植物性状提取中的鲁棒性、可扩展性与可靠性。

原文摘要 · Abstract (English)

Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and specialized retrieval models to improve ranking accuracy, relevance, and cross-lingual performance. Complementary techniques such as passage indexing, document layout analysis, and semantic knowledge representation further enhance retrieval effectiveness by capturing fine-grained contextual and structural information. Emerging agentic LLM frameworks extend these capabilities by enabling planning, iterative reasoning, tool use, and multi-agent collaboration, thereby broadening applications across diverse domains. These frameworks also emphasize rigorous evaluation, ethical considerations, and trustworthiness, ensuring responsible deployment in real-world settings. We propose a modular, agent-based pipeline for botanical trait extraction. Optical character recognition (OCR) converts PDFs into machine-readable text, while segmentation and indexing organize content by genus and species. Rule-based parsers extract structured botanical traits, and ensembles of large language models (LLMs) expand trait vocabularies and resolve ambiguities. This approach ensures accurate species recognition, scalable annotation, and explainable integration of textual botanical descriptions, enabling robust and interpretable data extraction across large botanical corpora. Results: Using three regional botanical datasets, our system extracted 55,737 trait annotations across 4,961 species, averaging 9.1 traits per species. Integration of LLM-based enrichment improved coverage for 75% of traits, increasing total annotations by 59%. While the choice of OCR engine had a minor effect on species recognition, overall annotation counts remained stable, demonstrating the robustness, scalability, and reliability of the pipeline for large-scale botanical trait extraction.

植物性状大模型文档解析知识抽取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。