小型语言模型可实现与大模型相当的医疗预测能力,且更省资源、更私密。
HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring
- 在零样本、少样本和指令微调下评估小语言模型性能
- 小模型在移动端部署后表现接近大模型,效率显著提升
- 适合注重隐私与实时性的可穿戴健康监测场景
移动与可穿戴健康监测对及时干预、慢病管理及提升生活质量至关重要。以往大语言模型(LLMs)在医疗预测任务中展现强大泛化能力,但多数基于云端,存在隐私风险并导致内存占用高、延迟大。为此,轻量级的小语言模型(SLMs)因其可在本地高效运行而受到关注。本文系统评估了SLMs在健康预测任务中的表现,采用零样本、少样本及指令微调策略,并将最优微调后的模型部署于移动设备,实测其在真实医疗场景下的效率与预测性能。结果表明,SLMs在性能上可媲美LLMs,同时大幅提高效率与隐私保护水平。然而,在类别不平衡与少样本场景中仍存挑战。当前形态虽不完美,但为下一代隐私友好型健康监测提供了可行路径。
原文摘要 · Abstract (English)
Mobile and wearable healthcare monitoring play a vital role in facilitating timely interventions, managing chronic health conditions, and ultimately improving individuals' quality of life. Previous studies on large language models (LLMs) have highlighted their impressive generalization abilities and effectiveness in healthcare prediction tasks. However, most LLM-based healthcare solutions are cloud-based, which raises significant privacy concerns and results in increased memory usage and latency. To address these challenges, there is growing interest in compact models, Small Language Models (SLMs), which are lightweight and designed to run locally and efficiently on mobile and wearable devices. Nevertheless, how well these models perform in healthcare prediction remains largely unexplored. We systematically evaluated SLMs on health prediction tasks using zero-shot, few-shot, and instruction fine-tuning approaches, and deployed the best performing fine-tuned SLMs on mobile devices to evaluate their real-world efficiency and predictive performance in practical healthcare scenarios. Our results show that SLMs can achieve performance comparable to LLMs while offering substantial gains in efficiency and privacy. However, challenges remain, particularly in handling class imbalance and few-shot scenarios. These findings highlight SLMs, though imperfect in their current form, as a promising solution for next-generation, privacy-preserving healthcare monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。