提升知识图谱生成文本质量,筛选出更可靠的三元组
Validating DBpedia Triple Sets for Natural Language Generation

- 基于自然语言生成视角,设计三元组筛选规则
- 三元组选择精度达98%,召回率提升40%且不降低精度
- 适合需要高质量知识源的文本生成研究者
我们从自然语言生成的角度研究了DBpedia三元组的质量,提出并评估了一种收集实体特定三元组集合的方法,该方法能过滤可疑三元组,同时最小化正确三元组的损失。在与人工标注数据对比的评估中,使用验证规则可实现98%的三元组选择精确率;通过改进少数属性定义,可在不损害精确率的前提下将召回率提升40%。
原文摘要 · Abstract (English)
We present a study of the quality of individual DBpedia triples from the perspective of Natural Language Generation, and propose and evaluate an approach for collecting entity-specific triple sets that filters out questionable triples while minimizing the loss of correct ones. We show in an evaluation against manually annotated data that with validation rules, it is possible to reach 98% precision in triple selection, and with improvements to a few Property definitions, it is possible to improve recall by 40% without harming precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。