arXiv:2607.21685cs.CL2026-07

评估方式影响专家与自动标注的差距,设计不同结果差异大

Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark

  • 用词袋和BiomedBERT对比专家与自动MeSH标注效果
  • 标准评估下专家标注优势明显,但换评估方式后差距缩小甚至消失
  • 研究提醒:论文结论可能受评估设计左右,适合方法论审慎者

系统性综述需筛选数千篇摘要,常借助分类器优先排序。分类器输入常加入医学主题词(MeSH),由专家数周后手动标注或自动工具即时生成。本文首次直接比较两者作为特征的性能,并探究评估设计对结果的影响。基于Cohen等(2006)药物分类基准,在三种主题上对比专家标注与基于子串匹配的机械式标注,分别使用词袋逻辑回归(7次重复)和BiomedBERT(5个种子)模型。在标准5折全语料设计下,词袋模型对他汀类药物的差距为+0.096 WSS@95%;通过分层抽样匹配语料规模(n=803)后,差距降至+0.033,置信区间包含零;10折交叉验证下进一步降至+0.021。BiomedBERT在标准设计下得+0.020,与词袋10折结果仅差0.001。单次运行的实证功效分析显示,其他两类主题在当前设计下无法检测到类似他汀类效应(奥美拉唑MDE=0.254,ADHD MDE=0.384);多轮重复协议下的有效样本量未被设计确定。结果仅适用于所测试的词汇匹配器,不推广至所有自动MeSH标注。总体而言,基准结论可能因评估设计改变而显著不同。

原文摘要 · Abstract (English)

A systematic review begins with someone reading thousands of abstracts to identify the few that are relevant, and classifiers are used to prioritise that reading. Their inputs are often augmented with Medical Subject Headings (MeSH), assigned either by expert indexers weeks or months after publication or by automatic tools at once. We did not identify prior work comparing the two directly as classifier features, or asking whether that comparison's outcome depends on how the classifier is evaluated. Using the Cohen et al. (2006) drug-class benchmark, we compare expert assignment against one mechanical procedure, substring matching against a MeSH vocabulary drawn from the benchmark, across a bag-of-words logistic regression classifier (seven reruns) and BiomedBERT (five seeds) on three topics. Under the canonical 5-fold full-corpus design the bag-of-words gap on Statins is +0.096 WSS@95%. Stratified subsampling to matched corpus size (n=803) reduces it by roughly two thirds, to +0.033, with a bootstrap interval that includes zero; 10-fold cross-validation at full corpus size reduces it by roughly four fifths, to +0.021. BiomedBERT under canonical evaluation gives +0.020, a difference of 0.001 from the bag-of-words 10-fold result. An empirical power analysis on a single canonical run per topic indicates that a Statins-sized effect at the per-fold variances of the other two topics would not have been detectable at that design (MDE 0.254 for Opioids, 0.384 for ADHD); at the pooled fold count of the multi-run protocol the bound depends on an effective sample size the design does not determine. The results bound the specific lexical matcher tested rather than automatic MeSH indexing in general. More broadly, benchmark conclusions about feature sources can change substantially under reasonable changes to the evaluation design.

信息检索评估设计自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。