现有指代消解评估存在有效性问题,模型表现依赖测试条件。
The Validity of Coreference-based Evaluations of Natural Language Understanding
- 通过扩展评估方法,检验指代判断中事件合理性的推理能力
- 当前模型在标准任务上表现优于基线,但泛化能力不足
- 研究揭示评估设计缺陷,适合关注评测可信度的研究者
本文通过拓展现有评估实践,深入分析指代消解评估的有效性问题。首先指出标准评估因定义争议和结果不一致(收敛性差)导致结论难以推广。其次提出一种新评估方法,聚焦系统对事件相对合理性的推断能力。结果显示,当代语言模型在标准基准上表现优于早期基线,在特定领域和类型上准确率提升,但在评估条件微调后普遍出现性能下降,无法像人类一样稳定泛化。这些发现既肯定了模型在主流评估中的进步,也暴露了当前NLP范式在测量有效性上的根本局限,为未来开发更可靠评估方法和更具泛化能力的系统提供了方向。
原文摘要 · Abstract (English)
In this thesis, I refine our understanding as to what conclusions we can reach from coreference-based evaluations by expanding existing evaluation practices and considering the extent to which evaluation results are either converging or conflicting. First, I analyze standard coreference evaluations and show that their design often leads to non-generalizable conclusions due to issues of measurement validity - including contestedness (multiple, competing definitions of coreference) and convergent validity (evaluation results that rank models differently across benchmarks). Second, I propose and implement a novel evaluation focused on testing systems' ability to infer the relative plausibility of events, a key aspect of resolving coreference. Through this extended evaluation, I find that contemporary language models demonstrate strong performance on standard benchmarks - improving over earlier baseline systems within certain domains and types of coreference - but remain sensitive to the evaluation conditions: they often fail to generalize in ways one would expect a human to be capable of when evaluation contexts are slightly modified. Taken together, these findings clarify both the strengths, such as improved accuracy over baselines on widely used evaluations, and the limitations of the current NLP paradigm, including weaknesses in measurement validity, and suggest directions for future work in developing better evaluation methods and more genuinely generalizable systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。