构建首个蛋白质互作生物效应评估基准,助力药物靶点发现
RAGPPI: RAG Benchmark for Protein-Protein Interactions in Drug Discovery
- 基于专家访谈设计问答对,聚焦蛋白互作的生物学影响
- 包含4420组问答,含500组人工标注黄金标准数据
- 提供自动评估框架,适合医药AI研究者使用
蛋白质-蛋白质互作(PPI)的生物学影响检索对药物研发中的靶点识别至关重要。由于涉及蛋白质数量庞大,该过程耗时且复杂。尽管大语言模型(LLMs)和检索增强生成(RAG)框架已支持靶点识别,但目前尚无专门用于评估PPI生物学影响的基准。为此,我们提出RAGPPI:一个包含4,420组问答对的事实性问题回答基准,聚焦于PPI的潜在生物学效应。通过与专家访谈,确立了问答类型和数据来源等基准标准。通过专家驱动的数据标注构建了500组黄金标准数据集。开发了一个集成专家标注特征、平均事实摘要相似度(F1)及低相似度事实数量(F2)的集成自动评估大语言模型,从而构建出3,720组银标准数据集。我们致力于将RAGPPI持续维护为支持药物发现领域RAG系统研究的公共资源。
原文摘要 · Abstract (English)
Retrieving the biological impacts of protein-protein interactions (PPIs) is essential for target identification (Target ID) in drug development. Given the vast number of proteins involved, this process remains time-consuming and challenging. Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) frameworks have supported Target ID; however, no benchmark currently exists for identifying the biological impacts of PPIs. To bridge this gap, we introduce the RAG Benchmark for PPIs (RAGPPI), a factual question-answer benchmark of 4,420 question-answer pairs that focus on the potential biological impacts of PPIs. Through interviews with experts, we identified criteria for a benchmark dataset, such as a type of QA and source. We built a gold-standard dataset (500 QA pairs) through expert-driven data annotation. We developed an ensemble auto-evaluation LLM that incorporates expert labeling characteristics, average fact-abstract similarity (F1), and low-similarity fact counts (F2), enabling the construction of a silver-standard dataset (3,720 QA pairs). We are committed to maintaining RAGPPI as a resource to support the research community in advancing RAG systems for drug discovery QA solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。