arXiv:2601.16711cs.CLcs.IR2026-01Conference of the …被引 1

用大模型自动生成标注数据,提升医学概念识别对未知概念的泛化能力。

Better Generalizing to Unseen Concepts: An Evaluation Framework and An LLM-Based Auto-Labeled Pipeline for Biomedical Concept Recognition

  • 构建基于层级概念索引的评估框架,量化模型对未知概念的泛化性能。
  • 利用大模型生成自标注数据,显著提升模型对未见概念的识别能力。
  • 适合关注医学信息抽取与小样本学习的研究者参考。

由于提及无关的生物医学概念识别(MA-BCR)中人工标注稀缺,对未见概念的泛化成为核心挑战。本文提出两个关键贡献:首先,构建基于层级概念索引的评估框架及新指标,系统衡量泛化性能;其次,探索基于大语言模型的自标注数据(ALD)作为可扩展资源,设计任务特定的生成管道。研究明确表明,尽管LLM生成的ALD无法完全替代人工标注,但其作为补充资源能有效提升模型泛化能力,提供更广覆盖与结构化知识,助力模型接近识别未见概念。代码与数据集见https://github.com/bio-ie-tool/hi-ald。

原文摘要 · Abstract (English)

Generalization to unseen concepts is a central challenge due to the scarcity of human annotations in Mention-agnostic Biomedical Concept Recognition (MA-BCR). This work makes two key contributions to systematically address this issue. First, we propose an evaluation framework built on hierarchical concept indices and novel metrics to measure generalization. Second, we explore LLM-based Auto-Labeled Data (ALD) as a scalable resource, creating a task-specific pipeline for its generation. Our research unequivocally shows that while LLM-generated ALD cannot fully substitute for manual annotations, it is a valuable resource for improving generalization, successfully providing models with the broader coverage and structural knowledge needed to approach recognizing unseen concepts. Code and datasets are available at https://github.com/bio-ie-tool/hi-ald.

医学信息抽取大模型应用概念识别泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。