arXiv:2511.21285cs.CL2025-11Conference of the …被引 5

评测7种高效微调方法在27个NLP任务上的表现,兼顾训练与推理成本。

PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark

  • 构建统一基准,覆盖7种参数高效微调方法
  • 在27个数据集上验证性能,支持可复现评估
  • 引入综合成本指标,平衡参数量、速度与内存

尽管大语言模型在诸多任务中表现出色,但其庞大的规模常带来高昂的计算与环境成本,限制了其可及性。参数高效微调(PEFT)方法通过减少可训练参数数量,在保持下游性能的同时缓解这一问题。然而现有评估仍受限于模型与数据集范围,且难以复现。为此,我们提出PEFT-Bench,一个面向自回归大语言模型的统一端到端基准,用于评估多种PEFT方法。我们在27个NLP数据集和7种PEFT方法上展示了其应用。为考量不同训练与推理因素,我们还引入了PEFT软成本惩罚(PSCP)指标,综合考虑可训练参数量、推理速度与训练内存占用。

原文摘要 · Abstract (English)

Despite the state-of-the-art performance of Large Language Models (LLMs) achieved on many tasks, their massive scale often leads to high computational and environmental costs, limiting their accessibility. Parameter-Efficient Fine-Tuning (PEFT) methods address this challenge by reducing the number of trainable parameters while maintaining strong downstream performance. Despite the advances in PEFT methods, current evaluations remain limited (in terms of evaluated models and datasets) and difficult to reproduce. To bridge this gap, we introduce PEFT-Bench, a unified end-to-end benchmark for evaluating diverse PEFT methods on autoregressive LLMs. We demonstrate its usage across 27 NLP datasets and 7 PEFT methods. To account for different PEFT training and inference factors, we also introduce the PEFT Soft Cost Penalties (PSCP) metric, which takes trainable parameters, inference speed, and training memory usage into account.

微调大模型评估基准效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。