为自然语料中的命题重要性设计可量化的标注方法。
Not Worth Mentioning? A Pilot Study on Salient Proposition Annotation
- 基于分级摘要思想定义命题重要性标注任务。
- 在多文体小数据集上验证标注一致性,初步发现与话语单元中心性的关联。
- 适合关注文本重要性评估与话语分析的研究者。
尽管提取式摘要研究已有长期传统,旨在恢复文本中最重要命题,但对自然语料中命题重要性的分级量化仍缺乏工作。本文采用先前命题实体抽取(SEE)研究中的分级摘要式重要性作为度量,并将其适配用于量化命题重要性。我们定义了标注任务,应用于一个小规模多文体数据集,评估了标注者间一致性,并初步探讨了该度量与基于修辞结构理论(RST)的话语单元中心性之间的关系。
原文摘要 · Abstract (English)
Despite a long tradition of work on extractive summarization, which by nature aims to recover the most important propositions in a text, little work has been done on operationalizing graded proposition salience in naturally occurring data. In this paper, we adopt graded summarization-based salience as a metric from previous work on Salient Entity Extraction (SEE) and adapt it to quantify proposition salience. We define the annotation task, apply it to a small multi-genre dataset, evaluate agreement and carry out a preliminary study of the relationship between our metric and notions of discourse unit centrality in discourse parsing following Rhetorical Structure Theory (RST).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。