首个多语言生物医学检索增强模型筛选能力评测基准
CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
- 构建多语言生物医学检索增强模型筛选能力评测基准
- 发现主流大模型在生物医学信息筛选上表现差异显著
- 适合关注生物医学AI评估与模型优化的研究者使用
检索增强型大语言模型在生物医学领域展现出巨大潜力,但其信息筛选能力的可靠评估仍存在关键空白。为此,我们提出首个面向生物医学检索增强模型筛选能力的多语言评测基准CRAB,涵盖英语、法语、德语和中文。通过引入基于引用的新评价指标,CRAB可量化评估模型在生物医学场景中的信息筛选性能。实验结果揭示主流大模型在该任务中存在显著性能差异,凸显提升其在生物医学领域筛选能力的紧迫性。数据集已公开于https://huggingface.co/datasets/zhm0/CRAB。
原文摘要 · Abstract (English)
Recent development in Retrieval-Augmented Large Language Models (LLMs) have shown great promise in biomedical applications. How ever, a critical gap persists in reliably evaluating their curation ability the process by which models select and integrate relevant references while filtering out noise. To address this, we introduce the benchmark for Curation of Retrieval-Augmented LLMs in Biomedicine (CRAB), the first multilingual benchmark tailored for evaluating the biomedical curation of retrieval-augmented LLMs, available in English, French, German and Chinese. By incorporating a novel citation-based evaluation metric, CRAB quantifies the curation performance of retrieval-augmented LLMs in biomedicine. Experimental results reveal significant discrepancies in the curation performance of mainstream LLMs, underscoring the urgent need to improve it in the domain of biomedicine. Our dataset is available at https://huggingface.co/datasets/zhm0/CRAB.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。