arXiv:2608.03529cs.CL2026-08

提出量化非结构化生物医学标注一致性的新方法

Consensus Measures for Unstructured Biomedical Text Annotations

论文配图:Consensus Measures for Unstructured Biomedical Text Annotations
图 1 · 摘自论文原文
  • 用语义等价度量评估开放标签的标注者间一致性
  • 嵌入模型可扩展但难区分相似概念,LLM易受规模限制
  • 推荐基于自然语言推理的度量作为平衡方案

生物医学文献正被用于挖掘其撰写初衷之外的知识。由于目标概念事先未知,标注者更倾向使用开放式文本标签,这使得一致性难以量化。本文研究了在提供非结构化文本的生物医学标注任务中,软性标注者间可靠性(soft inter-rater reliability)的度量方法。合成实验表明,多种语义等价度量可用于量化软性可靠性,且度量选择会影响估计的失败模式。嵌入模型具有可扩展性,但在区分相似但不同的概念时表现有限;大型语言模型虽有潜力,但因估算随机一致性而面临可扩展性瓶颈。最终建议采用基于自然语言推理的度量作为合理折衷方案。

原文摘要 · Abstract (English)

Biomedical literature is increasingly mined for knowledge beyond the questions it was written to answer. Because the target concepts are not known in advance, annotators prefer open-ended labels, whose agreement is hard to quantify. We study soft inter-rater reliability for annotators providing unstructured texts for biomedical annotation tasks. Synthetic experiments show that soft reliability can be quantified using a variety of semantic equivalence measures, and that the choice of measure affects failure modes of the estimation. Embeddings are scalable, but limited when differentiating similar but distinct concepts. Large language models are promising, but limited by scalability for estimating agreement by chance. Finally, we suggest measures based on natural language inference as a sensible compromise.

标注一致性生物医学自然语言推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。