arXiv:2503.17425cs.CLcs.IR2025-03中稿 · Text2Story Worksho…被引 3

针对临床文本中的断言检测,构建了更精准的模型,超越通用大模型表现。

Beyond Negation Detection: Comprehensive Assertion Detection Models for Clinical NLP

  • 基于微调大模型和深度学习,提升断言分类精度
  • 在存在、缺失、假设等类别上显著优于GPT-4o与商业API
  • 轻量少样本模型适合资源受限场景,可集成到医疗信息处理流程

断言状态识别是临床自然语言处理中关键但常被忽视的环节,对准确归属医学事实至关重要。以往研究多局限于否定检测,导致AWS Medical Comprehend、Azure AI Text Analytics及GPT-4o等商业方案因领域适应性差而表现不佳。本文开发了先进的断言检测模型,包括微调大模型、基于Transformer的分类器、少样本分类器及深度学习方法。评估显示,微调大模型整体准确率达0.962,显著优于GPT-4o(0.901)和商业API;在‘存在’(+4.2%)、‘缺失’(+8.4%)、‘假设’(+23.4%)类别上优势明显。深度学习模型在‘条件’(+5.3%)和‘他人关联’(+10.1%)类别超越商业方案,少样本模型达0.929,适用于资源受限环境。集成于Spark NLP后,模型持续优于黑箱商业方案,支持可扩展推理与医学命名实体识别、关系抽取、术语解析无缝衔接。结果表明,领域适配、透明可控的临床NLP方案优于通用大模型与专有API。

原文摘要 · Abstract (English)

Assertion status detection is a critical yet often overlooked component of clinical NLP, essential for accurately attributing extracted medical facts. Past studies have narrowly focused on negation detection, leading to underperforming commercial solutions such as AWS Medical Comprehend, Azure AI Text Analytics, and GPT-4o due to their limited domain adaptation. To address this gap, we developed state-of-the-art assertion detection models, including fine-tuned LLMs, transformer-based classifiers, few-shot classifiers, and deep learning (DL) approaches. We evaluated these models against cloud-based commercial API solutions, the legacy rule-based NegEx approach, and GPT-4o. Our fine-tuned LLM achieves the highest overall accuracy (0.962), outperforming GPT-4o (0.901) and commercial APIs by a notable margin, particularly excelling in Present (+4.2%), Absent (+8.4%), and Hypothetical (+23.4%) assertions. Our DL-based models surpass commercial solutions in Conditional (+5.3%) and Associated-with-Someone-Else (+10.1%) categories, while the few-shot classifier offers a lightweight yet highly competitive alternative (0.929), making it ideal for resource-constrained environments. Integrated within Spark NLP, our models consistently outperform black-box commercial solutions while enabling scalable inference and seamless integration with medical NER, Relation Extraction, and Terminology Resolution. These results reinforce the importance of domain-adapted, transparent, and customizable clinical NLP solutions over general-purpose LLMs and proprietary APIs.

临床NLP断言检测大模型微调医疗信息提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。