对比大模型性能,提出医疗领域选型方法
Mirror, Mirror on the Wall -- Which is the Best Model of Them All?
- 构建系统性选型框架,结合榜单数据与任务需求
- 以医学领域为例,揭示主流模型性能演进趋势
- 适合需要精准选型的开发者和研究者参考
大型语言模型(LLMs)在金融、医疗、教育、通信、法律等多个领域显著提升了生产力并取得优异成果。当前主流基础模型由大公司基于海量数据和高昂算力训练而成,但新模型迭代迅速,选择合适模型愈发复杂。本文聚焦量化评估维度,分析现有排行榜与基准测试,以医学领域为案例,展示模型性能演变现状及实际意义。在此基础上,提出模型选型方法(MSM),提供一套系统化流程,帮助用户根据具体应用场景,有效导航、优先排序并选择最优模型。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have become one of the most transformative tools across many applications, as they have significantly boosted productivity and achieved impressive results in various domains such as finance, healthcare, education, telecommunications, and law, among others. Typically, state-of-the-art (SOTA) foundation models are developed by large corporations based on large data collections and substantial computational and financial resources required to pretrain such models from scratch. These foundation models then serve as the basis for further development and domain adaptation for specific use cases or tasks. However, given the dynamic and fast-paced nature of launching new foundation models, the process of selecting the most suitable model for a particular use case, application, or domain becomes increasingly complex. We argue that there are two main dimensions that need to be taken into consideration when selecting a model for further training: a qualitative dimension (which model is best suited for a task based on information, for instance, taken from model cards) and a quantitative dimension (which is the best performing model). The quantitative performance of models is assessed through leaderboards, which rank models based on standardized benchmarks and provide a consistent framework for comparing different LLMs. In this work, we address the analysis of the quantitative dimension by exploring the current leaderboards and benchmarks. To illustrate this analysis, we focus on the medical domain as a case study, demonstrating the evolution, current landscape, and practical significance of this quantitative evaluation dimension. Finally, we propose a Model Selection Methodology (MSM), a systematic approach designed to guide the navigation, prioritization, and selection of the model that best aligns with a given use case.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。