arXiv:2411.09944cs.CL2024-11被引 15

SlimLM让小模型在手机上高效处理文档,兼顾速度与隐私。

SlimLM: An Efficient Small Language Model for On-Device Document Assistance

  • 针对手机优化小模型,平衡参数量、上下文长度与推理时间。
  • 最小模型在三星S24上运行高效,大模型在移动端仍具更强能力。
  • 提供安卓应用与基准数据集,适合移动AI开发与隐私敏感场景。

尽管小型语言模型(SLMs)在移动端部署中展现潜力,但其真实性能与应用场景仍不明确。本文提出SlimLM,一系列专为手机端文档辅助任务优化的小型语言模型。通过在三星Galaxy S24上的大量实验,我们确定了模型规模(125M至7B参数)、上下文长度与推理时间之间的最佳权衡。SlimLM在SlimPajama-627B上预训练,并在自建的DocAssist数据集上微调,涵盖摘要、问答与建议任务。最小模型在S24上表现高效,更大版本在移动约束下仍具备更强能力。与现有SLMs对比,SlimLM表现相当或更优,并为未来设备端语言模型研究提供基准。我们还发布了安卓应用,提供实际部署洞察。研究结果表明,高阶智能手机可有效运行先进语言模型,有望降低服务器成本并提升隐私保护。

原文摘要 · Abstract (English)

While small language models (SLMs) show promises for mobile deployment, their real-world performance and applications on smartphones remains underexplored. We present SlimLM, a series of SLMs optimized for document assistance tasks on mobile devices. Through extensive experiments on a Samsung Galaxy S24, we identify the optimal trade-offs between model size (ranging from 125M to 7B parameters), context length, and inference time for efficient on-device processing. SlimLM is pre-trained on SlimPajama-627B and fine-tuned on DocAssist, our constructed dataset for summarization, question answering and suggestion tasks. Our smallest model demonstrates efficient performance on S24, while larger variants offer enhanced capabilities within mobile constraints. We evaluate SlimLM against existing SLMs, showing comparable or superior performance and offering a benchmark for future research in on-device language models. We also provide an Android application, offering practical insights into SLM deployment. Our findings provide valuable insights and illuminate the capabilities of running advanced language models on high-end smartphones, potentially reducing server costs and enhancing privacy through on-device processing.

小模型手机端文档辅助隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。