语言模型在维诺格拉德难题上表现好,却解决不了更简单的指代消解任务。
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
- 用提示词语言模型结合监督式系统提升指代消解性能
- 提示模型在维诺格拉德难题上得分高,但在OntoNotes上表现差
- 单一数据集评估无法反映系统真实能力,需多任务验证
维诺格拉德难题(WSC)等挑战集被用来评估系统处理自然语言歧义的能力。现有研究通常假设,若模型在挑战集上表现优异,则其在更通用任务上也应表现良好。然而我们通过实证发现,这一假设并不总是成立。具体而言,尽管提示式语言模型(LMs)在WSC及其变体上表现强劲,但其在OntoNotes等数据集中看似更简单的指代消解任务上表现相对较差。基于此,我们提出一种方法:将提示式语言模型与一个专用于指代消解的监督式系统进行集成,以提升跨数据集的整体准确性。最后,我们强调,涉及同一语言现象的不同数据集依赖于不同的、但有重叠的能力,仅在一个数据集上评估无法全面反映系统的综合能力。
原文摘要 · Abstract (English)
Challenge sets such as the Winograd Schema Challenge (WSC) are used to benchmark systems' ability to resolve ambiguities in natural language. If one assumes as in existing work that solving a given challenge set is at least as difficult as solving some more general task, then high performance on the challenge set should indicate high performance on the general task overall. However, we show empirically that this assumption of difficulty does not always hold. In particular, we demonstrate that despite the strong performance of prompted language models (LMs) on the WSC and its variants, these same modeling techniques perform relatively poorly at resolving certain pronominal ambiguities attested in OntoNotes and related datasets that are perceived to be easier. Motivated by these findings, we propose a method for ensembling a prompted LM with a supervised, task-specific system that is overall more accurate at resolving pronominal coreference across datasets. Finally, we emphasize that datasets involving the same linguistic phenomenon draw on distinct, but overlapping, capabilities, and evaluating on any one dataset alone does not provide a complete picture of a system's overall capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。