arXiv:2509.01190cs.CL2025-09

无需微调即可动态加速大模型推理,最高提速11倍。

Efficient Large Language Models with Zero-Shot Adjustable Acceleration

  • 推理时动态调整硬件利用,不需额外微调。
  • 在多个任务上实现最高11倍加速,性能稳定。
  • 适合追求高效部署的工程师和研究者。

在真实应用场景中使用大语言模型(LLMs)面临计算效率与模型性能平衡的重大挑战。优化微调后及推理过程中的加速能力对构建高效架构至关重要。本文提出零样本可调节加速(Zero-Shot Adjustable Acceleration),一种新颖的训练与推理方法,可在推理阶段动态调整硬件利用率,无需额外微调。该方法应用于近期主流大语言模型,并在多个分类与文本生成任务上进行评估。实验结果表明,该方法支持广泛的零样本加速,相比基线模型最高可达11倍速度提升。

原文摘要 · Abstract (English)

Using Large Language Models (LLMs) in real-world applications presents significant challenges, particularly in balancing computational efficiency with model performance. Optimizing acceleration after fine-tuning and during inference is critical for building efficient architectures. This paper introduces Zero-Shot Adjustable Acceleration, a novel training and inference method that dynamically adjusts hardware utilization during inference without requiring additional fine-tuning. The proposed approach is applied to recent LLMs and evaluated across multiple classification and text generation tasks. Experimental results demonstrate that the method supports a wide range of zero-shot acceleration and achieves up to 11x speedup compared to the baseline.

大模型加速推理优化零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。