arXiv:2606.00994cs.CL2026-06

用四重机制构建可审计的物种性状提取系统,实现大规模精准记录。

A Registry-Bound LLM Pipeline for Evidence-Grounded Trait Extraction across Tropical Plants, Aquatic Species, and Exotic Pets

论文配图:A Registry-Bound LLM Pipeline for Evidence-Grounded Trait Extraction across Tropical Plants, Aquatic Species, and Exotic Pets
图 1 · 摘自论文原文
  • 基于版本化性状注册表约束输出,确保结构规范
  • 每条记录附原文引用与置信度标签,99.985%物种成功提取
  • 三重验证保障证据可信,适合生物数据标准化需求

我们提出一种基于注册表的大型语言模型提取流水线,可规模化生成栽培热带植物、水生生物和异宠物种的证据支撑型结构化性状记录。四个机制保障输出可审计:版本化的39个键值闭合词汇性状注册表,将每个取值限定于类型化模式;每条记录附原文引文,关联原始文本;每条记录标注置信度(高或中;低者不持久化);多版本保存。在《热带物种百科全书》409,880个可发表物种上执行706,220次运行,共持久化5,489,881条性状记录,覆盖409,820个物种(99.985%),其中81.57%为高置信度。报告三重验证层:全量样本中,5,427,588条带证据行中有90.12%的引文是源文本的完整子串(剔除一个合规元性状后为93.49%);对100条分层非红区样本的引文-值一致性审计得100/100(下限96.30%);对50条红区样本的外观有效性评估得50/50接受(下限92.86%)。单条记录正确性未声称,仍待人工校验。贡献在于四机制框架。

原文摘要 · Abstract (English)

We describe a registry-bound large-language-model extraction pipeline producing evidence-grounded structured trait records at scale, on cultivated tropical plant, aquatic, and pet species. Four mechanisms render LLM-derived rows auditable: a versioned 39-key closed-vocabulary trait registry constraining every admitted value to a typed schema; a per-row verbatim evidence quote tying each value to source text; a per-row confidence label (high or medium; low dropped pre-persist); and multi-version preservation. Applied to 409,880 publishable species from the Tropical Species Encyclopedia, the pipeline executed 706,220 runs and persisted 5,489,881 trait records across 409,820 species (99.985%), 81.57% at high confidence. We report three validation layers in descending evidentiary strength: at full population, 90.12% of 5,427,588 evidence-bearing rows have their quote as a verbatim source substring (93.49% excluding one compliance meta-trait); a quote-supports-value audit on n=100 stratified non-red-zone rows yielded 100/100 (lower bound 96.30%); face-validity on n=50 red-zone rows yielded 50/50 Accept (lower bound 92.86%). Per-record correctness is not claimed; 100% pending human curation. The contribution is the four-mechanism framework.

知识图谱生物信息大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。