arXiv:2509.07512cs.CLcs.AI2025-09EMNLP被引 2

用三阶段主动学习精选演示样本,降低大模型实体识别标注成本

ALLabel: Three-stage Active Learning for LLM-based Entity Recognition using Demonstration Retrieval

  • 分三阶段筛选最有信息量的标注样本用于大模型提示学习
  • 仅标注5%-10%数据即可达到全量标注效果
  • 适合需要高效标注科学领域实体的团队使用

当前自然科学领域的许多数据驱动研究,如化学和材料科学,都需要从科学数据集中进行大规模、高性能的实体识别。大语言模型(LLMs)在该任务中应用日益广泛,与全谱自然语言处理任务趋势一致。现有基于大模型的实体识别方法多依赖微调,但该过程成本高昂。为实现性能与成本的最佳平衡,本文提出ALLabel,一种三阶段框架,通过选择最具信息量和代表性样本构建大模型提示学习所需的演示语料库。该框架依次采用三种不同的主动学习策略,在三个专业领域数据集上均以相同标注预算优于所有基线方法。实验表明,仅需标注5%-10%的数据,即可达到全量标注的性能水平。进一步分析与消融实验证明了该方法的有效性与通用性。

原文摘要 · Abstract (English)

Many contemporary data-driven research efforts in the natural sciences, such as chemistry and materials science, require large-scale, high-performance entity recognition from scientific datasets. Large language models (LLMs) have increasingly been adopted to solve the entity recognition task, with the same trend being observed on all-spectrum NLP tasks. The prevailing entity recognition LLMs rely on fine-tuned technology, yet the fine-tuning process often incurs significant cost. To achieve a best performance-cost trade-off, we propose ALLabel, a three-stage framework designed to select the most informative and representative samples in preparing the demonstrations for LLM modeling. The annotated examples are used to construct a ground-truth retrieval corpus for LLM in-context learning. By sequentially employing three distinct active learning strategies, ALLabel consistently outperforms all baselines under the same annotation budget across three specialized domain datasets. Experimental results also demonstrate that selectively annotating only 5\%-10\% of the dataset with ALLabel can achieve performance comparable to the method annotating the entire dataset. Further analyses and ablation studies verify the effectiveness and generalizability of our proposal.

主动学习实体识别大模型标注效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。