arXiv:2512.07068cs.CL2025-12被引 1

提出首个英文到统一语义表示的解析模型,提升低资源语言技术潜力。

SETUP: Sentence-level English-To-Uniform Meaning Representation Parser

  • 基于抽象语义表示与通用依存关系转换,构建双路径解析框架。
  • 在AnCast上达84分,SMATCH++达91分,显著优于已有基线。
  • 适合语义解析、跨语言研究及低资源语言技术开发者参考。

统一语义表示(UMR)是一种基于图结构的新型语义表示,能够捕捉文本核心含义,并通过灵活的标注模式支持全球多种语言(包括低资源语言)的标注。尽管UMR在语言记录、低资源语言技术改进和可解释性方面展现出潜力,但其下游应用仍受限于缺乏高效的文本到UMR解析器。目前关于文本到UMR解析的研究十分有限。本文提出两种英文文本到UMR的解析方法:一种是微调现有的抽象语义表示(AMR)解析器,另一种则利用通用依存关系(Universal Dependencies)转换器作为基础。我们提出的最佳模型SETUP,在AnCast评分中达到84分,SMATCH++评分为91分,表明在自动化生成准确的UMR图方面取得了显著进展。

原文摘要 · Abstract (English)

Uniform Meaning Representation (UMR) is a novel graph-based semantic representation which captures the core meaning of a text, with flexibility incorporated into the annotation schema such that the breadth of the world's languages can be annotated (including low-resource languages). While UMR shows promise in enabling language documentation, improving low-resource language technologies, and adding interpretability, the downstream applications of UMR can only be fully explored when text-to-UMR parsers enable the automatic large-scale production of accurate UMR graphs at test time. Prior work on text-to-UMR parsing is limited to date. In this paper, we introduce two methods for English text-to-UMR parsing, one of which fine-tunes existing parsers for Abstract Meaning Representation and the other, which leverages a converter from Universal Dependencies, using prior work as a baseline. Our best-performing model, which we call SETUP, achieves an AnCast score of 84 and a SMATCH++ score of 91, indicating substantial gains towards automatic UMR parsing.

语义解析UMR自然语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。