构建首个英文语料库,为无时态信息的语义图补全事件时态标签。
Enhancing Structured Meaning Representations with Aspect Classification
- 基于AMR图构建带时态标签的标注体系,统一事件时态分类标准。
- 通过多轮审校流程确保标注一致性,建立高质量数据集。
- 提供基线模型结果,推动自动时态信息预测研究进展。
为完整捕捉句子语义,语义表示需包含描述事件内部时间结构的时态信息。在基于图的语义表示框架(如UMR)中,时态可揭示事件随时间展开的方式,包括状态、活动与完成事件等区分。尽管重要,现有语义表示框架中时态标注仍极为稀疏,阻碍了人工标注及自动系统的发展。本文构建了一个新的英文语料库,对缺乏时态特征的抽象语义表示(AMR)图进行UMR时态标签标注。我们详述了依据UMR时态层级对谓词进行标注的方案与指南,并通过多步仲裁流程确保标注者间的一致性与质量。为展示该数据集对未来自动化任务的价值,我们采用三种建模范式开展基线实验。结果确立了自动UMR时态预测的初始基准,为更广泛集成时态信息于语义表示提供了基础。
原文摘要 · Abstract (English)
To fully capture the meaning of a sentence, semantic representations should encode aspect, which describes the internal temporal structure of events. In graph-based meaning representation frameworks such as Uniform Meaning Representations (UMR), aspect lets one know how events unfold over time, including distinctions such as states, activities, and completed events. Despite its importance, aspect remains sparsely annotated across semantic meaning representation frameworks. This has, in turn, hindered not only current manual annotation, but also the development of automatic systems capable of predicting aspectual information. In this paper, we introduce a new dataset of English sentences annotated with UMR aspect labels over Abstract Meaning Representation (AMR) graphs that lack the feature. We describe the annotation scheme and guidelines used to label eventive predicates according to the UMR aspect lattice, as well as the annotation pipeline used to ensure consistency and quality across annotators through a multi-step adjudication process. To demonstrate the utility of our dataset for future automation, we present baseline experiments using three modeling approaches. Our results establish initial benchmarks for automatic UMR aspect prediction and provide a foundation for integrating aspect into semantic meaning representations more broadly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。