arXiv:2608.06167cs.AIcs.CL2026-08

用生成式AI从文本中提取复杂结构信息并自动评估准确性

Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI

  • 基于模式的框架一次性零样本提取嵌套属性
  • 14个属性中12个提取F1超90%,耗时仅为人工1/30
  • 支持多模型、多机构、多语言迁移,适用于医疗评估领域

我们提出一种基于模式的框架,利用生成式AI从非结构化文本中提取复杂的结构化信息,并对提取结果与标准答案进行自动化语义评估。该模式作为编码领域知识的信息模型,为层级嵌套、变基数属性的提取与评估提供统一系统框架。信息提取通过一次模型调用完成,无需微调。评估阶段引入路径匹配算法,对提取结果与标准答案中的嵌套属性进行对齐,并使用生成式AI比较属性值语义,依据领域标准分类为精确、语义一致、有用或不匹配。在健康技术评估组织NICE发布的文档上,使用Claude Opus 3模型成功提取14个属性中的12个,F1分数超过90%;单文档提取时间约为人类专家的1/30。该框架还展现出跨不同生成式AI模型、不同HTA组织及多种语言的泛化与迁移能力。

原文摘要 · Abstract (English)

We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard. The schema, serving as an information model encoding domain knowledge, provides a unified, systematic, and consistent framework for extraction of hierarchical, nested information, with attributes of variable cardinality, and subsequent evaluation of the results. Information extraction from a document is performed in a single call to the model, in zero-shot mode. In the evaluation step, we introduce a path-based semantic matching algorithm to align the nested, variable-cardinality attributes in the extracted results with those in the gold standard. We use generative AI for semantic comparison of the extracted and gold standard values of an attribute, and introduce a rubric to classify the result of the comparison, according to domain-specific considerations, as an exact, semantic, useful, or non-match. We were able to extract 12 out of 14 attributes with an F1 score of $>$90\% from documents published by the health technology assessment organisation NICE, using the generative AI model Claude Opus 3. The time needed to extract the attributes from a document was $\sim$30 times lower than the time taken by a human domain expert. We further demonstrate generalisability of this framework across different generative AI models and transferability across different HTA organisations and languages.

信息抽取生成式AI医疗评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。