arXiv:2603.00042cs.LGcs.AI2026-03

通过对齐潜在空间几何,实现亚1比特大模型压缩的性能突破。

LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment

  • 引入内部潜空间旋转与联合迭代量化,优化二值化分布
  • 在0.1~1比特每参数下超越现有1比特方法,匹配主流1比特模型精度
  • 零推理开销,适合部署资源受限的大模型场景

我们发现,在极端模型压缩中存在谱能量增益现象:对于重尾谱分布,低秩二值近似优于微小秩浮点基线。然而,以往方法未能发挥此潜力,性能落后于领先的1比特方法。我们归因于潜在空间几何错位——标准奇异向量具有高相干性(尖峰分布),是二值量化最不利的几何结构。为此,提出LittleBit-2框架,采用内部潜空间旋转与联合迭代量化(Joint-ITQ),作为几何预处理器,将相干潜空间分布对齐至二值超立方体,且不增加推理开销。实验证明,LittleBit-2在Llama-2和Llama-3上实现亚1比特(1~0.1 bpp)领域的全新最优性能,其精度可媲美领先的1比特基线方法。

原文摘要 · Abstract (English)

We identify the Spectral Energy Gain in extreme model compression, where low-rank binary approximations outperform tiny-rank floating-point baselines for heavy-tailed spectra. However, prior attempts fail to realize this potential, trailing state-of-the-art 1-bit methods. We attribute this degradation to Latent Geometry Misalignment: standard singular vectors exhibit high coherence (spiky distribution), the worst-case geometry for binary quantization. To realize this gain, we propose LittleBit-2, a framework employing Internal Latent Rotation and Joint Iterative Quantization (Joint-ITQ). This approach acts as a geometric preconditioner, aligning coherent latent distributions with the binary hypercube with zero inference overhead. Empirically, LittleBit-2 establishes a new state-of-the-art in the sub-1-bit regime (1$\sim$0.1 bpp) on Llama-2 and Llama-3, matching the fidelity of leading 1-bit baselines.

大模型压缩二值量化潜空间对齐高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。