arXiv:2604.13232cs.CL2026-04

指出语义变化检测基准存在三方面缺陷,提醒勿将其视为绝对标准。

Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection

  • 将语义变化简化为词义增减,忽略渐进与语境变化
  • 数据含大量排版错误、标注偏差,影响模型表现与可复现性
  • 小规模目标集+有限语言覆盖,导致结果不够真实可信

本文通过操作化、数据质量与基准设计三个层面重新审视SemEval-2020 Task 1这一最具影响力的词汇语义变化检测共享基准。首先,该任务将语义变化主要定义为离散词义的增、减或重构,但忽略了渐进式、构式性、搭配性及话语层变化;此外,黄金标签受标注决策、聚类过程与阈值设定影响,可能削弱任务有效性。其次,数据存在严重问题:包括OCR噪声、乱码字符、句子截断、不一致的词形还原、词性标注错误及目标词遗漏,这些会扭曲模型行为,阻碍语言分析并降低可复现性。第三,任务采用小规模精选目标词集且语言覆盖有限,导致评估设定缺乏现实性,并加剧统计不确定性。综上,该基准应被视为有用但非完备的测试平台。我们呼吁未来数据集与评测任务应采纳更广泛的语义变化理论,透明披露文档预处理流程,扩展跨语言覆盖,并采用更真实的评估设置,以实现更具有效性、可解释性和泛化性的研究进展。

原文摘要 · Abstract (English)

This discussion paper re-examines SemEval-2020 Task 1, the most influential shared benchmark for lexical semantic change detection, through a three-part evaluative framework: operationalisation, data quality, and benchmark design. First, at the level of operationalisation, we argue that the benchmark models semantic change mainly as gain, loss, or redistribution of discrete senses. While practical for annotation and evaluation, this framing is too narrow to capture gradual, constructional, collocational, and discourse-level change. Also, the gold labels are outcomes of annotation decisions, clustering procedures, and threshold settings, which could potentially limit the validity of the task. Second, at the level of data quality, we show that the benchmark is affected by substantial corpus and preprocessing problems, including OCR noise, malformed characters, truncated sentences, inconsistent lemmatisation, POS-tagging errors, and missed targets. These issues can distort model behaviour, complicate linguistic analysis, and reduce reproducibility. Third, at the level of bench-mark design, we argue the small curated target sets and limited language coverage reduce realism and increase statistical uncertainty. Taken together, these limitations suggest that the benchmark should be treated as a useful but partial test bed rather than a definitive measure of progress. We therefore call for future datasets and shared tasks to adopt broader theories of semantic change, document pre-processing transparently, expand cross-linguistic coverage, and use more realistic evaluation settings. Such steps are necessary for more valid, interpretable, and generalisable progress in lexical semantic change detection

语义变化评估基准数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。