arXiv:2511.12472cs.CLcs.AI2025-11

用新框架评估大模型在药物重定位中发现意外洞见的能力

Assessing LLMs for Serendipity Discovery in Knowledge Graphs: A Case for Drug Repurposing

  • 构建了基于相关性、新颖性和惊喜度的惊异性评估指标
  • 大模型在知识检索上表现良好,但发现真正意外答案仍困难
  • 适合关注科学发现创新性与大模型探索能力的研究者

大语言模型(LLMs)在知识图谱问答(KGQA)方面取得显著进展,但现有系统通常优化为返回高度相关且可预测的答案。本文提出一种新的惊异性感知KGQA任务,并设计了SerenQA框架,用于评估LLMs在科学KGQA任务中揭示意外洞察的能力。该框架包含基于相关性、新颖性和惊喜度的严谨惊异性度量,以及基于临床知识图谱(Clinical Knowledge Graph)的专家标注基准,聚焦药物重定位。此外,还设计了包含知识检索、子图推理和惊异性探索三个子任务的结构化评估流程。实验表明,尽管先进LLMs在检索任务上表现优异,但在识别真正令人惊喜且有价值的发现方面仍存在明显不足,凸显未来改进空间。相关资源已开源:https://cwru-db-group.github.io/serenQA。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have greatly advanced knowledge graph question answering (KGQA), yet existing systems are typically optimized for returning highly relevant but predictable answers. A missing yet desired capacity is to exploit LLMs to suggest surprise and novel ("serendipitious") answers. In this paper, we formally define the serendipity-aware KGQA task and propose the SerenQA framework to evaluate LLMs' ability to uncover unexpected insights in scientific KGQA tasks. SerenQA includes a rigorous serendipity metric based on relevance, novelty, and surprise, along with an expert-annotated benchmark derived from the Clinical Knowledge Graph, focused on drug repurposing. Additionally, it features a structured evaluation pipeline encompassing three subtasks: knowledge retrieval, subgraph reasoning, and serendipity exploration. Our experiments reveal that while state-of-the-art LLMs perform well on retrieval, they still struggle to identify genuinely surprising and valuable discoveries, underscoring a significant room for future improvements. Our curated resources and extended version are released at: https://cwru-db-group.github.io/serenQA.

知识图谱药物重定位大模型评估惊异性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。