20亿参数的1比特大模型,性能媲美全精度模型。
BitNet b1.58 2B4T Technical Report
- 采用原生1比特权重设计,大幅降低计算开销。
- 4万亿词训练,性能达同类开源模型水平。
- 支持多架构部署,适合资源受限场景使用。
我们提出BitNet b1.58 2B4T,首个在20亿参数规模上开源的原生1比特大型语言模型。该模型在4万亿词语料上训练,并在涵盖语言理解、数学推理、编程能力及对话表现的多个基准上进行了严格评估。结果表明,BitNet b1.58 2B4T在性能上与同规模领先的开源全精度模型相当,同时在计算效率方面具有显著优势,包括大幅降低内存占用、能耗和解码延迟。为促进后续研究与应用,模型权重已通过Hugging Face发布,并提供支持GPU与CPU架构的开源推理实现。
原文摘要 · Abstract (English)
We introduce BitNet b1.58 2B4T, the first open-source, native 1-bit Large Language Model (LLM) at the 2-billion parameter scale. Trained on a corpus of 4 trillion tokens, the model has been rigorously evaluated across benchmarks covering language understanding, mathematical reasoning, coding proficiency, and conversational ability. Our results demonstrate that BitNet b1.58 2B4T achieves performance on par with leading open-weight, full-precision LLMs of similar size, while offering significant advantages in computational efficiency, including substantially reduced memory footprint, energy consumption, and decoding latency. To facilitate further research and adoption, the model weights are released via Hugging Face along with open-source inference implementations for both GPU and CPU architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。