测试手机端大模型在医疗推理中的表现,发现小型通用模型兼顾速度与准确。
Medicine on the Edge: Comparative Performance Analysis of On-Device LLMs for Clinical Reasoning
- 用AMEGA数据集对比手机上运行的多个大模型性能。
- 小模型Phi-3 Mini速度准度平衡好,医学微调模型准确率最高。
- 旧手机也能用,内存比算力更关键,适合临床场景落地。
将大语言模型(LLM)部署于移动设备可显著提升医疗应用的隐私性、安全性与成本效益,避免依赖云端服务并使敏感健康数据本地化。然而,现有研究对手机端LLM在真实医疗场景下的表现仍缺乏系统评估。本研究基于AMEGA数据集,对公开可用的手机端LLM进行基准测试,评估其在准确性、计算效率及热限制方面的表现。结果表明,紧凑型通用模型如Phi-3 Mini在速度与准确率间取得良好平衡,而医学领域微调模型如Med42和Aloe达到最高准确率。值得注意的是,即使在较旧设备上部署也具备可行性,但内存限制比算力瓶颈更具挑战性。研究证实了手机端LLM在医疗领域的潜力,同时强调需开发更高效的推理机制及面向真实临床推理的专用模型。
原文摘要 · Abstract (English)
The deployment of Large Language Models (LLM) on mobile devices offers significant potential for medical applications, enhancing privacy, security, and cost-efficiency by eliminating reliance on cloud-based services and keeping sensitive health data local. However, the performance and accuracy of on-device LLMs in real-world medical contexts remain underexplored. In this study, we benchmark publicly available on-device LLMs using the AMEGA dataset, evaluating accuracy, computational efficiency, and thermal limitation across various mobile devices. Our results indicate that compact general-purpose models like Phi-3 Mini achieve a strong balance between speed and accuracy, while medically fine-tuned models such as Med42 and Aloe attain the highest accuracy. Notably, deploying LLMs on older devices remains feasible, with memory constraints posing a greater challenge than raw processing power. Our study underscores the potential of on-device LLMs for healthcare while emphasizing the need for more efficient inference and models tailored to real-world clinical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。