用类型和实体信息提升长尾关系抽取准确率,比大模型更高效。
DISCIE -- Discriminative Closed Information Extraction
- 基于类型与实体的判别式方法,聚焦长尾关系。
- 在百万实体、数百关系场景下超越主流生成模型。
- 小模型实现大模型性能,适合大规模高效抽取。
本文提出一种新型封闭式信息抽取方法,采用判别式框架,融合类型与实体特异性信息,显著提升关系抽取准确率,尤其对长尾关系效果突出。在包含数百万实体与数百关系的大规模场景下,该方法表现优于当前最先进的端到端生成模型。通过使用小型模型实现高效计算,证明引入类型信息可使小模型性能达到甚至超过大型生成模型水平,为更精准、高效的抽取技术提供新路径。
原文摘要 · Abstract (English)
This paper introduces a novel method for closed information extraction. The method employs a discriminative approach that incorporates type and entity-specific information to improve relation extraction accuracy, particularly benefiting long-tail relations. Notably, this method demonstrates superior performance compared to state-of-the-art end-to-end generative models. This is especially evident for the problem of large-scale closed information extraction where we are confronted with millions of entities and hundreds of relations. Furthermore, we emphasize the efficiency aspect by leveraging smaller models. In particular, the integration of type-information proves instrumental in achieving performance levels on par with or surpassing those of a larger generative model. This advancement holds promise for more accurate and efficient information extraction techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。