用结构化框架从医学PDF中自动提取可审计的证据,提升文献综述效率与可信度。
From Chaos to Clarity: Schema-Constrained AI for Auditable Biomedical Evidence Extraction from Full-Text PDFs
- 通过类型约束的模式和证据门控决策,限制模型推理以提高准确性
- 在抗凝药水平研究数据上实现零人工干预、稳定吞吐与高一致性提取
- 支持溯源与事后审计,适合需要透明可靠的医学证据合成场景
生物医学证据合成依赖于从全文科研论文中准确提取方法学、实验及结果变量,但这些信息深藏于复杂的科学PDF中,手动抽取耗时且难以扩展。现有文档AI系统受限于OCR错误、长文档碎片化、处理吞吐量不足以及高风险合成场景下的审计性缺失。我们提出一种基于模式约束的AI提取系统,通过显式使用类型化模式、受控词汇表和证据门控决策,将全文字医学PDF转化为结构化、可分析的记录。文档采用感知恢复的哈希技术摄入,按含图注的页面分块,并在显式并发控制下异步处理。块级输出通过冲突感知合并、集合聚合与句子级溯源机制,确定性地整合为研究级记录,支持可追溯性与事后审计。在直接口服抗凝药水平测量的研究语料库上评估,该流程实现了全程无人工干预、服务约束下稳定吞吐,且各文档块间保持强内部一致性。迭代模式优化显著提升了关键变量(如检测方法分类、结局定义、随访时长与测量时间)的提取精度。结果表明,模式约束与溯源感知的提取方法可实现异构科学PDF向结构化证据的可扩展、可审计转化,使现代文档AI更契合生物医学证据合成对透明性与可靠性的要求。
原文摘要 · Abstract (English)
Biomedical evidence synthesis relies on accurate extraction of methodological, laboratory, and outcome variables from full-text research articles, yet these variables are embedded in complex scientific PDFs that make manual abstraction time-consuming and difficult to scale. Existing document AI systems remain limited by OCR errors, long-document fragmentation, constrained throughput, and insufficient auditability for high-stakes synthesis. We present a schema-constrained AI extraction system that transforms full-text biomedical PDFs into structured, analysis-ready records by explicitly restricting model inference through typed schemas, controlled vocabularies, and evidence-gated decisions. Documents are ingested using resume-aware hashing, partitioned into caption-aware page-level chunks, and processed asynchronously under explicit concurrency controls. Chunk-level outputs are deterministically merged into study-level records using conflict-aware consolidation, set-based aggregation, and sentence-level provenance to support traceability and post-hoc audit. Evaluated on a corpus of studies on direct oral anticoagulant level measurement, the pipeline processed all documents without manual intervention, maintained stable throughput under service constraints, and exhibited strong internal consistency across document chunks. Iterative schema refinement substantially improved extraction fidelity for synthesis-critical variables, including assay classification, outcome definitions, follow-up duration, and timing of measurement. These results demonstrate that schema-constrained, provenance-aware extraction enables scalable and auditable transformation of heterogeneous scientific PDFs into structured evidence, aligning modern document AI with the transparency and reliability requirements of biomedical evidence synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。