arXiv:2602.09469cs.CLcs.AI2026-02

用集成模型精准识别西班牙语病历中的药物使用与上下文信息

NOWJ @BioCreative IX ToxHabits: An Ensemble Deep Learning Approach for Detecting Substance Use and Contextual Information in Clinical Texts

  • 融合BETO与CRF的多输出集成模型,提升序列标注效果
  • 触发词检测达0.94 F1、0.97精确率,论元检测0.91 F1
  • 适合低资源临床文本分析,对可解释性要求高的场景

从非结构化电子健康记录中提取药物使用信息仍是临床自然语言处理的重大挑战。尽管大语言模型取得进展,但在临床NLP中的应用受限于信任、控制和效率问题。为此,我们提出参与BioCreative IX ToxHabits共享任务的NOWJ方案。该任务聚焦于西班牙语临床文本中成瘾物质使用的检测及其上下文属性识别,属于领域特定、低资源场景。我们构建了一个多输出集成系统,同时应对子任务1(ToxNER)和子任务2(ToxUse)。系统结合BETO与CRF层进行序列标注,采用多样训练策略,并通过句子过滤提升精确率。最佳结果在触发词检测上达到0.94 F1和0.97精确率,论元检测达0.91 F1。

原文摘要 · Abstract (English)

Extracting drug use information from unstructured Electronic Health Records remains a major challenge in clinical Natural Language Processing. While Large Language Models demonstrate advancements, their use in clinical NLP is limited by concerns over trust, control, and efficiency. To address this, we present NOWJ submission to the ToxHabits Shared Task at BioCreative IX. This task targets the detection of toxic substance use and contextual attributes in Spanish clinical texts, a domain-specific, low-resource setting. We propose a multi-output ensemble system tackling both Subtask 1 - ToxNER and Subtask 2 - ToxUse. Our system integrates BETO with a CRF layer for sequence labeling, employs diverse training strategies, and uses sentence filtering to boost precision. Our top run achieved 0.94 F1 and 0.97 precision for Trigger Detection, and 0.91 F1 for Argument Detection.

临床NLP药物使用检测多任务学习低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。