arXiv:2606.19345cs.CLcs.AI2026-06

用大模型集成自动识别医学文献中的健康生活质量数据。

Ensembles of Large Language Models for Identifying EQ-5D Studies in PubMed Based on Their Abstracts

  • 融合多个大模型的少样本提示与加权集成策略。
  • 集成模型达到0.74的加权F1分数,优于单个模型。
  • 适合需要高效筛选医学文献的研究人员使用。

科研论文数量激增导致系统性文献综述中人工筛选越来越耗时、低效且不一致。识别明确报告健康相关生活质量结果(如EQ-5D数据)的研究,需较高临床判断力,对人工评审构成挑战。本研究探讨基于仅含摘要的PubMed文献,利用Google的Gemini和Gemma大语言模型自动检测EQ-5D研究的可行性。提出多阶段框架,包含少样本提示、权重集成聚合及软堆叠元分类器。在两名专家手动标注的PubMed研究数据集上评估了九个LLM。由gemini-2.5-pro、gemma-3-12b和gemma-3-27b组成的加权集成模型获得0.74的加权F1分数和0.74准确率,优于各模型单独表现。集成方法提升了精确率与召回率的平衡,软堆叠策略增强了可靠性与可解释性。特征分析显示,模型输出的概率值对最终预测具有重要引导作用。结果表明,基于集成的大模型方案是自动化生物医学研究筛选的可靠且可扩展的方法。

原文摘要 · Abstract (English)

The rapid increase in scientific publications leads to the fact that manual study screening in systematic literature reviews (SLRs) is increasingly resource consuming, inefficient, and inconsistent. Classifying studies that clearly report health-related quality-of-life results, such as EQ-5D data, requires a high level of clinical interpretation and poses challenges for human reviewers. This study investigates the use of Google's Gemini and Gemma large language models (LLMs) in automating EQ-5D detection in the PubMed biomedical database based only on published abstracts. A multi-phase framework is proposed that integrates few-shot prompting, weight ensembling aggregation, and a soft stacking meta-classifier. Nine LLMs are evaluated on a dataset of PubMed studies manually labeled by two experts regarding EQ-5D reporting. The weighted ensemble of gemini-2.5-pro, gemma-3-12b, and gemma-3-27b obtained a 0.74 weighted F1-score and 0.74 accuracy, exceeding individually attained results. The ensembling of top-performing models improved the balance between precision and recall compared to individual models, while the soft stacking approach provided greater reliability and interpretability. Feature analysis shows that the probability results from the models are important in guiding the final predictions. The findings suggest that an ensemble-based LLM setup is a reliable and scalable approach for automating screening in biomedical research.

大模型文献筛选健康质量集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。