arXiv:2601.06624cs.CL2026-01被引 1

用抽样方法高效估算生物医学实体链接质量,仅需24.6%标注量即可保证精度。

Efficient and Reliable Estimation of Named Entity Linking Quality: A Case Study on GutBrainIE

  • 基于分层两阶段聚类抽样,构建独立于标注结果的标签与表面形式分组。
  • 在GutBrainIE数据集上仅标注2749条三元组,误差控制在0.05以内,准确率91.5%±4.73%。
  • 相比随机抽样减少约29%标注时间,适合大规模生物医学信息抽取评估。

命名实体链接(NEL)是生物医学信息抽取的核心组件,但因专家标注成本高、语料规模大,难以实现规模化质量评估。本文提出一种基于抽样的框架,在有限标注预算下,以统计保证估算大规模信息抽取语料的NEL准确率。将准确率估计建模为约束优化问题,目标是最小化预期标注成本,同时满足指定的误差范围(MoE)。借鉴知识图谱准确性评估的研究,将分层两阶段聚类抽样(STWCS)适配至NEL场景,定义基于标签的分层与全局表面形式聚类,且不依赖于已有的NEL标注。应用于2025年秋季公开发布的新型生物医学语料GutBrainIE,该框架在手动标注2,749个三元组(占总数24.6%)的情况下,实现了≤0.05的误差范围,总体准确率估计为0.915±0.0473。时间成本模型和与简单随机抽样(SRS)基线的模拟表明,相同样本量下可节省约29%的专家标注时间。该框架具有通用性,适用于其他NEL基准与需可扩展、统计稳健评估的信息抽取流程。

原文摘要 · Abstract (English)

Named Entity Linking (NEL) is a core component of biomedical Information Extraction (IE) pipelines, yet assessing its quality at scale is challenging due to the high cost of expert annotations and the large size of corpora. In this paper, we present a sampling-based framework to estimate the NEL accuracy of large-scale IE corpora under statistical guarantees and constrained annotation budgets. We frame NEL accuracy estimation as a constrained optimization problem, where the objective is to minimize expected annotation cost subject to a target Margin of Error (MoE) for the corpus-level accuracy estimate. Building on recent works on knowledge graph accuracy estimation, we adapt Stratified Two-Stage Cluster Sampling (STWCS) to the NEL setting, defining label-based strata and global surface-form clusters in a way that is independent of NEL annotations. Applied to 11,184 NEL annotations in GutBrainIE -- a new biomedical corpus openly released in fall 2025 -- our framework reaches a MoE $\leq 0.05$ by manually annotating only 2,749 triples (24.6%), leading to an overall accuracy estimate of $0.915 \pm 0.0473$. A time-based cost model and simulations against a Simple Random Sampling (SRS) baseline show that our design reduces expert annotation time by about 29% at fixed sample size. The framework is generic and can be applied to other NEL benchmarks and IE pipelines that require scalable and statistically robust accuracy assessment.

实体链接生物医学抽样评估统计可靠

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。