arXiv:2506.05766cs.CL2025-06被引 5

构建多模态药物相互作用问答数据集,推动LLM跨模态推理能力发展

BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions

  • 基于文本与分子结构的多模态知识图谱支持信息检索
  • 现有LLM在无背景数据时准确率不足,依赖外部知识才表现良好
  • 适合研究生物医学多模态推理与RAG框架的学者使用

检索增强生成(RAG)在提升大语言模型(LLM)性能方面展现出巨大潜力。然而,现有基于RAG的LLM主要聚焦于单一模态信息检索,以文本为主;而现实问题如医疗领域中,相关信息可能以知识图谱、文本(临床记录)及复杂分子结构等多种模态呈现。因此,能够检索多模态特定领域信息,并综合多种知识进行推理以生成准确回答至关重要。为填补这一空白,我们提出BioMol-MQA,一个专注于多药联用(polypharmacy)的多模态问答数据集,包含两部分:(i) 包含文本和分子结构的多模态知识图谱(KG),用于信息检索;(ii) 设计用于测试LLM在多模态KG上检索与推理能力的挑战性问题。基准测试表明,现有LLM在未提供必要背景数据时难以正确回答这些问题,仅在获得外部信息支持后表现良好,凸显了强大RAG框架的必要性。

原文摘要 · Abstract (English)

Retrieval augmented generation (RAG) has shown great power in improving Large Language Models (LLMs). However, most existing RAG-based LLMs are dedicated to retrieving single modality information, mainly text; while for many real-world problems, such as healthcare, information relevant to queries can manifest in various modalities such as knowledge graph, text (clinical notes), and complex molecular structure. Thus, being able to retrieve relevant multi-modality domain-specific information, and reason and synthesize diverse knowledge to generate an accurate response is important. To address the gap, we present BioMol-MQA, a new question-answering (QA) dataset on polypharmacy, which is composed of two parts (i) a multimodal knowledge graph (KG) with text and molecular structure for information retrieval; and (ii) challenging questions that designed to test LLM capabilities in retrieving and reasoning over multimodal KG to answer questions. Our benchmarks indicate that existing LLMs struggle to answer these questions and do well only when given the necessary background data, signaling the necessity for strong RAG frameworks.

多模态知识图谱药物相互作用LLM推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。