arXiv:2601.03266cs.CLcs.AI2026-01

小模型本地运行也能精准辅助临床决策,还更隐私安全。

Benchmarking and Adapting On-Device LLMs for Clinical Decision Support

  • 在本地设备上测试多个开源大模型,对比诊断表现。
  • 微调后模型准确率最高达87.9%,接近商用模型水平。
  • 适合医疗场景对隐私和算力有限的机构使用。

大型语言模型(LLMs)在临床决策中进展迅速,但专有系统因隐私顾虑和依赖云端部署而受限。开源模型虽支持本地推理,但多数体积过大,难以在资源受限的临床环境中应用。本文评测了gpt-oss(20b、120b)、Qwen3.5(9B、27B、35B)和Gemma 4(31B)系列的本地运行模型,在三大典型临床任务上的表现:通用疾病诊断、眼科专科诊断与管理、以及模拟专家评分评估。对比了当前领先专有模型(GPT-5.1、GPT-5-mini、Gemini 3.1 Pro)及主流开源模型DeepSeek-R1。进一步对gpt-oss-20b和Qwen3.5-35B在通用诊断数据上进行微调。结果表明,本地模型性能可媲美或超越DeepSeek-R1和GPT-5-mini,且显著更小。微调后,Qwen3.5-35B诊断准确率达87.9%,接近专有模型GPT-5.1(89.4%)。基础模型中,Gemma 4 31B通用诊断准确率为86.5%,优于GPT-5-mini,并接近微调后的Qwen3.5-35B。错误分析显示,87.2%的诊断错误为临床合理鉴别诊断,而非无关回答;上限分析表明通过优化答案选择可达93.2%准确率。结果表明,本地运行的LLMs具备高准确性、强适应性与隐私保护优势,是推动其在临床实践中广泛应用的可行路径。

原文摘要 · Abstract (English)

Large language models (LLMs) have rapidly advanced in clinical decision-making, yet the deployment of proprietary systems is hindered by privacy concerns and reliance on cloud-based infrastructure. Open-source alternatives allow local inference but often have large model sizes that limit their use in resource-constrained clinical settings. Here, we benchmark on-device LLMs from the gpt-oss (20b, 120b), Qwen3.5 (9B, 27B, 35B), and Gemma 4 (31B) families across three representative clinical tasks: general disease diagnosis, specialty-specific (ophthalmology) diagnosis and management, and simulation of human expert grading and evaluation. We compare their performance with state-of-the-art proprietary models (GPT-5.1, GPT-5-mini, and Gemini 3.1 Pro) and a leading open-source model (DeepSeek-R1), and we further evaluate the adaptability of on-device systems by fine-tuning gpt-oss-20b and Qwen3.5-35B on general diagnostic data. Across tasks, on-device models achieve performance comparable to or exceeding DeepSeek-R1 and GPT-5-mini despite being substantially smaller. In addition, fine-tuning remarkably improves diagnostic accuracy, with the fine-tuned Qwen3.5-35B reaching 87.9% and approaching the proprietary GPT-5.1 (89.4%). Among base on-device models, Gemma 4 31B achieved the strongest general diagnostic accuracy at 86.5%, exceeding GPT-5-mini and approaching the fine-tuned Qwen3.5-35B. Error characterization revealed that 87.2% of diagnostic errors across all models were clinically plausible differentials rather than off-topic predictions, and upper-bound analysis showed up to 93.2% attainable accuracy through improved answer selection. These findings highlight the potential of on-device LLMs to deliver accurate, adaptable, and privacy-preserving clinical decision support, offering a practical pathway for broader integration of LLMs into routine clinical practice.

临床决策本地部署大模型医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。