arXiv:2411.17713cs.DCcs.AI2024-11被引 29

轻量版AI安全防护模型,手机端也能高效运行。

Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations

  • 采用量化压缩技术,模型仅440MB,体积缩小7倍。
  • 在安卓手机上实现每秒30词吞吐,首字延迟低于2.5秒。
  • 安全检测效果媲美甚至超越更大模型,适合移动端部署。

本文介绍了 Llama Guard 3-1B-INT4,一个轻量高效的 Llama Guard 模型,已在 Meta Connect 2024 开源。实验表明,该模型可部署于资源受限设备,在主流安卓手机的CPU上实现不低于每秒30个词的吞吐量,且首字生成时间不超过2.5秒。值得注意的是,尽管其大小仅为 440MB,约为 Llama Guard 3-1B 的七分之一,但其在安全内容过滤任务中的表现与原模型相当或更优。

原文摘要 · Abstract (English)

This paper presents Llama Guard 3-1B-INT4, a compact and efficient Llama Guard model, which has been open-sourced to the community during Meta Connect 2024. We demonstrate that Llama Guard 3-1B-INT4 can be deployed on resource-constrained devices, achieving a throughput of at least 30 tokens per second and a time-to-first-token of 2.5 seconds or less on a commodity Android mobile CPU. Notably, our experiments show that Llama Guard 3-1B-INT4 attains comparable or superior safety moderation scores to its larger counterpart, Llama Guard 3-1B, despite being approximately 7 times smaller in size (440MB).

AI安全轻量化移动端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。