arXiv:2508.00121cs.CL2025-08中稿 · 16th IWCS

神经语义解析器在省略句上表现不佳,原因在于语言上下文复杂。

Is neural semantic parsing good at ellipsis resolution, or isn't it?

  • 构建120个省略句的语义标注数据集,测试解析器表现
  • 标准测试集得分超90%,但省略句解析错误率显著上升
  • 省略句失败主因是上下文复杂,而非复制语义信息困难

神经语义解析器在多种语言现象上表现良好,语义匹配得分超过90%。然而,对于强依赖上下文、需重复大量语义信息才能形成完整表征的现象,其表现如何?以英语动词短语省略为例,整个动词短语可被单个助动词替代。我们构建了包含120个省略案例及其完整语义表示的语料库,作为对一批大型神经语义解析器的挑战任务。尽管这些解析器在标准测试集上表现优异,但在省略句上却出现明显失效。数据增强有助于提升性能。解析省略句困难的原因并非复制语义内容本身困难,而是其常出现在语言结构复杂的上下文中,导致多数解析错误。

原文摘要 · Abstract (English)

Neural semantic parsers have shown good overall performance for a variety of linguistic phenomena, reaching semantic matching scores of more than 90%. But how do such parsers perform on strongly context-sensitive phenomena, where large pieces of semantic information need to be duplicated to form a meaningful semantic representation? A case in point is English verb phrase ellipsis, a construct where entire verb phrases can be abbreviated by a single auxiliary verb. Are the otherwise known as powerful semantic parsers able to deal with ellipsis or aren't they? We constructed a corpus of 120 cases of ellipsis with their fully resolved meaning representation and used this as a challenge set for a large battery of neural semantic parsers. Although these parsers performed very well on the standard test set, they failed in the instances with ellipsis. Data augmentation helped improve the parsing results. The reason for the difficulty of parsing elided phrases is not that copying semantic material is hard, but that usually occur in linguistically complicated contexts causing most of the parsing errors.

语义解析省略句神经模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。