构建医学文本知识图谱复杂推理数据集,助力大模型精准医疗问答
RiTeK: A Dataset for Large Language Models Complex Reasoning over Textual Knowledge Graphs in Medicine
- 合成融合多种拓扑结构与关系信息的医疗查询
- 11种检索模型在该数据集上表现均不理想
- 为医疗领域大模型检索提供新基准,适合研究者使用
医学领域复杂现实问题的解答通常需要从医学文本知识图谱(medical TKGs)中准确检索信息,因为图谱中的关系路径能增强大语言模型(LLMs)的推理能力。然而,当前主要瓶颈在于现有医学TKGs数量稀少、拓扑结构表达力有限,且缺乏对现有检索器的全面评估。为此,我们构建了面向大模型医学文本知识图谱复杂推理的数据集RiTeK,涵盖广泛拓扑结构。具体地,我们合成包含多样化拓扑结构、关系信息和复杂文本描述的真实用户查询,并通过医学专家严格评估验证其质量。RiTeK同时作为评估基于LLM检索系统能力的综合性基准。对11种代表性检索器在该基准上的评估显示,现有方法表现不佳,揭示了当前基于LLM的检索方法在半结构化医学数据上的显著局限性。这凸显了开发更高效医学检索系统的需求。
原文摘要 · Abstract (English)
Answering complex real-world questions in the medical domain often requires accurate retrieval from medical Textual Knowledge Graphs (medical TKGs), as the relational path information from TKGs could enhance the inference ability of Large Language Models (LLMs). However, the main bottlenecks lie in the scarcity of existing medical TKGs, the limited expressiveness of their topological structures, and the lack of comprehensive evaluations of current retrievers for medical TKGs. To address these challenges, we first develop a Dataset1 for LLMs Complex Reasoning over medical Textual Knowledge Graphs (RiTeK), covering a broad range of topological structures. Specifically, we synthesize realistic user queries integrating diverse topological structures, relational information, and complex textual descriptions. We conduct a rigorous medical expert evaluation process to assess and validate the quality of our synthesized queries. RiTeK also serves as a comprehensive benchmark dataset for evaluating the capabilities of retrieval systems built upon LLMs. By assessing 11 representative retrievers on this benchmark, we observe that existing methods struggle to perform well, revealing notable limitations in current LLM-driven retrieval approaches. These findings highlight the pressing need for more effective retrieval systems tailored for semi-structured data in the medical domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。