多大模型协作提升医学问答准确率
Collaboration among Multiple Large Language Models for Medical Question Answering
- 设计框架让多个大模型协作答题
- 显著增强推理能力并减少答案分歧
- 适合医疗AI研究与临床辅助系统开发
得益于庞大的内部知识库,新一代大语言模型在处理医学任务方面展现出巨大潜力。然而,目前尚缺乏对多个大模型专业知识协同效应的充分探索。本研究提出一种针对医学多选题数据集的多大模型协作框架。通过对三个预训练大模型进行事后分析,该框架被证明能有效提升所有模型的推理能力,并缓解其在不同问题间的答案分歧。我们还测量了大模型在面对其他模型的对立意见时的置信度,发现其置信度与预测准确性存在一致关系。
原文摘要 · Abstract (English)
Empowered by vast internal knowledge reservoir, the new generation of large language models (LLMs) demonstrate untapped potential to tackle medical tasks. However, there is insufficient effort made towards summoning up a synergic effect from multiple LLMs' expertise and background. In this study, we propose a multi-LLM collaboration framework tailored on a medical multiple-choice questions dataset. Through post-hoc analysis on 3 pre-trained LLM participants, our framework is proved to boost all LLMs reasoning ability as well as alleviate their divergence among questions. We also measure an LLM's confidence when it confronts with adversary opinions from other LLMs and observe a concurrence between LLM's confidence and prediction accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。