arXiv:2508.18609cs.CLcs.AI2025-08ACL被引 1

提出细粒度知识量化规律,揭示不同能力对精度、规模、校准的敏感差异。

Task-Stratified Knowledge Scaling Laws for Post-Training Quantized Large Language Models

  • 按记忆、应用、推理分层构建量化规律框架
  • 低比特下校准集大小影响记忆能力,精度决定推理表现
  • 适用于设计更智能的低比特大模型部署策略

后训练量化(PTQ)是高效部署大语言模型的关键策略。现有缩放规律多关注整体性能,忽视了细粒度因素及量化对不同知识能力的差异化影响。为此,我们建立任务分层的知识缩放规律,将能力分为记忆、应用和推理三类,构建统一框架,融合模型规模、位宽、组大小和校准集大小等细粒度因素。在293种多样化的PTQ配置上验证,框架拟合良好且跨架构一致。结果表明:推理对精度敏感,应用对规模响应强,记忆对校准集大小敏感。在低比特场景下,优化这些细粒度因素对防止性能崩溃至关重要。研究为设计知识感知的量化策略提供了实证基础。

原文摘要 · Abstract (English)

Post-Training Quantization (PTQ) is a critical strategy for efficient Large Language Models (LLMs) deployment. However, existing scaling laws primarily focus on general performance, overlooking crucial fine-grained factors and how quantization differentially impacts diverse knowledge capabilities. To address this, we establish Task-Stratified Knowledge Scaling Laws. By stratifying capabilities into memorization, application, and reasoning, we develop a framework that unifies model size, bit-width, and fine-grained factors: group size and calibration set size. Validated on 293 diverse PTQ configurations, our framework demonstrates strong fit and cross-architecture consistency. It reveals distinct sensitivities across knowledge capabilities: reasoning is precision-critical, application is scale-responsive, and memorization is calibration-sensitive. We highlight that in low-bit scenarios, optimizing these fine-grained factors is essential for preventing performance collapse. These findings provide an empirically-backed foundation for designing knowledge-aware quantization strategies.

量化大模型缩放定律知识能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。