arXiv:2605.30826cs.CLcs.AI2026-05被引 2

用多模型共识优化生物医学实体候选排序,提升人工校对效率。

Beyond Agreement: Scoring Panel-Surfaced Biomedical Entity Candidates for Curator Triage

论文配图:Beyond Agreement: Scoring Panel-Surfaced Biomedical Entity Candidates for Curator Triage
图 1 · 摘自论文原文
  • 构建多模型联合输出的候选评测基准,以候选条目为单位验证。
  • 在域内测试中,模型准确率从0.753提升至0.910,可选1340个高精度候选。
  • 适合需高效筛选实体候选的生物医学文本标注团队使用。

现代大语言模型虽易识别生物医学术语,但其结果是否符合语料规范取决于标注惯例、边界定义、粒度和类型体系。单一模型预测的共识不能代表语料规范正确性。本文提出一种基于多模型面板输出的候选级评测基准,将八个大模型在五个公开生物医学命名实体识别数据集上的预测对齐成候选主表。BioConCal是一种域内监督评分器,利用推理时无需真实标签的共识、提及特征、表面可见性及文档特征,对固定候选流进行评分。在域内测试中,其AUROC从原始共识的0.753提升至0.910;在设定0.95精确率目标下,可选出1,340个候选,实测精确率达0.939,远超原始共识的293个;对应候选级召回率为0.592,语料级召回率为0.523(面板行标签上限为0.883)。主要优势并非找回所有模型均遗漏的实体,而是将噪声候选流重构为更高效的审阅队列。当实体类型分布发生变化时,阈值需在目标域验证,而精确字符定位仍需独立后处理步骤。

原文摘要 · Abstract (English)

Biomedical NER is deceptively simple for modern LLMs: plausible biomedical mentions are easy to surface, but corpus-convention correctness depends on annotation conventions, span boundaries, entity granularity, and type schemas. Multi-LLM agreement is a salience signal, not corpus-convention correctness. We introduce a candidate-level panel-output benchmark for panel-surfaced candidate verification, where the unit is an aligned candidate surfaced by an explicitly defined multi-model panel rather than a standalone extractor output. The benchmark aligns eight LLMs' predictions over five public biomedical NER datasets into a candidate master table. BioConCal is an in-domain supervised scorer that instantiates this layer with inference-time gold-free agreement, mention, surface-availability, and document features for a fixed candidate stream. In domain, BioConCal improves AUROC from 0.753 for raw agreement to 0.910. At a validation-selected 0.95 precision target it selects 1,340 candidates at empirical test precision 0.939, compared with 293 for raw agreement. This corresponds to candidate-level recall 0.592 and corpus-level recall 0.523 against a within-panel row-label ceiling of 0.883. The main benefit is not recovering entities missed by every panel member, but reshaping a noisy panel stream into a higher-yield review queue. Under entity-type shift, thresholds require target-domain validation, and exact character localization remains a separate deterministic post-processing step.

生物医学实体识别多模型融合标注效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。