提出新型多布尔架构,实现大模型直接布尔域微调,大幅降低计算复杂度。
Highly Efficient and Effective LLMs with Multi-Boolean Architectures
- 用多核布尔参数替代传统权重,实现布尔域直接微调
- 在多个LLM上超越最新超低比特量化与二值化方法
- 无需高精度中间权重,显著提升训练与推理效率
权重二值化已成为降低大语言模型(LLMs)复杂度的有前景策略。现有方法分为后训练二值化(简单但性能损失严重)和训练感知方法(依赖全精度潜在权重,增加复杂性并限制效率)。本文提出一种新框架,以多核布尔参数表示LLMs,首次实现布尔域直接微调,无需潜在权重。该方法增强表征能力,并在微调与推理阶段显著降低复杂度。在多种LLM上的大量实验表明,本方法优于近期超低比特量化与二值化技术。
原文摘要 · Abstract (English)
Weight binarization has emerged as a promising strategy to reduce the complexity of large language models (LLMs). Existing approaches fall into post-training binarization, which is simple but causes severe performance loss, and training-aware methods, which depend on full-precision latent weights, adding complexity and limiting efficiency. We propose a novel framework that represents LLMs with multi-kernel Boolean parameters and, for the first time, enables direct finetuning LMMs in the Boolean domain, eliminating the need for latent weights. This enhances representational capacity and dramatically reduces complexity during both finetuning and inference. Extensive experiments across diverse LLMs show our method outperforms recent ultra low-bit quantization and binarization techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。