arXiv:2601.09730cs.CLcs.AI2026-01综述

梳理临床文档元数据提取研究,揭示从规则到大模型的演进路径。

Clinical Document Metadata Extraction: A Scoping Review

  • 基于规则与机器学习转向基于Transformer和大模型的方法
  • 公开标注数据稀缺,仅结构段落数据较丰富
  • 适合医疗文本处理、电子病历系统研发者参考

临床文档元数据(如文档类型、结构、作者角色、医学专科和就诊场景)对准确解读临床信息至关重要。然而,文档异质性及随时间演变的差异给元数据标准化带来挑战。自动化提取方法应运而生,用于将不同实践中的元数据整合到目标模式中。本范围综述旨在梳理临床文档元数据提取研究,识别方法趋势与应用类型,并指出研究空白。我们遵循PRISMA-ScR指南,筛选2011年1月至2025年8月间发表的266篇文献,最终纳入67篇相关研究。其中45篇为方法学研究,17篇将元数据作为下游任务特征,5篇分析元数据构成。研究目的多样,方法已从依赖人工特征工程的规则与传统机器学习,发展为少特征工程的Transformer架构。大语言模型的兴起拓展了跨任务与数据集的泛化能力,推动更先进的临床文本处理系统发展。未来研究有望深化元数据表征并融入临床应用与工作流程。

原文摘要 · Abstract (English)

Clinical document metadata, such as document type, structure, author role, medical specialty, and encounter setting, is essential for accurate interpretation of information captured in clinical documents. However, vast documentation heterogeneity and drift over time challenge harmonization of document metadata. Automated extraction methods have emerged to coalesce metadata from disparate practices into target schema. This scoping review aims to catalog research on clinical document metadata extraction, identify methodological trends and applications, and highlight gaps. We followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines to identify articles that perform clinical document metadata extraction. We initially found and screened 266 articles published between January 2011 and August 2025, then comprehensively reviewed 67 we deemed relevant to our study. Among the articles included, 45 were methodological, 17 used document metadata as features in a downstream application, and 5 analyzed document metadata composition. We observe myriad purposes for methodological study and application types. Available labelled public data remains sparse except for structural section datasets. Methods for extracting document metadata have progressed from largely rule-based and traditional machine learning with ample feature engineering to transformer-based architectures with minimal feature engineering. The emergence of large language models has enabled broader exploration of generalizability across tasks and datasets, allowing the possibility of advanced clinical text processing systems. We anticipate that research will continue to expand into richer document metadata representations and integrate further into clinical applications and workflows.

元数据提取医疗文本大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。