用文献生成定义,提升科学文本跨文档指代与层级关系识别
Inferring Scientific Cross-Document Coreference and Hierarchy with Definition-Augmented Relational Reasoning
- 通过检索全文文献生成概念的上下文定义,增强指代识别
- 在表面形式多样、歧义高的数据上性能显著提升,尤其在少样本场景
- 适合知识图谱构建、学术搜索与发现等需要细粒度理解的场景
我们解决科学文本中跨文档指代与层级关系推断这一基础任务,该任务对知识图谱构建、搜索、推荐与发现具有重要意义。大语言模型在面对大量长尾技术概念及其细微差异时表现不佳。本文提出一种新方法:通过检索全文文献为概念提及生成上下文相关的定义,并利用这些定义增强跨文档关系的检测能力。进一步生成描述两个概念提及之间关系或差异的关联定义,并设计高效重排序策略,应对跨论文链接推断中的组合爆炸问题。在微调和上下文学习两种设置下,我们在表面形式多样且高度模糊的数据子集上均取得显著性能提升。我们还对生成的定义进行分析,揭示了大语言模型在细粒度科学概念上的关系推理能力。
原文摘要 · Abstract (English)
We address the fundamental task of inferring cross-document coreference and hierarchy in scientific texts, which has important applications in knowledge graph construction, search, recommendation and discovery. Large Language Models (LLMs) can struggle when faced with many long-tail technical concepts with nuanced variations. We present a novel method which generates context-dependent definitions of concept mentions by retrieving full-text literature, and uses the definitions to enhance detection of cross-document relations. We further generate relational definitions, which describe how two concept mentions are related or different, and design an efficient re-ranking approach to address the combinatorial explosion involved in inferring links across papers. In both fine-tuning and in-context learning settings, we achieve large gains in performance on data subsets with high amount of different surfaces forms and ambiguity, that are challenging for models. We provide analysis of generated definitions, shedding light on the relational reasoning ability of LLMs over fine-grained scientific concepts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。