提出小模型也能通过工具和推理实现大模型能力,重新定义压缩目标。
The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?
- 主张压缩时应保留多步推理与外部工具调用能力
- 指出当前压缩方法忽视了关键高级能力的保持
- 适合关注模型轻量化与功能保留的研究者
为降低大语言模型(LLM)的计算与存储开销,模型压缩和键值缓存(KV cache)压缩受到广泛关注。然而,现有方法主要关注压缩后模型在困惑度或常识问答、基础算术推理等任务上的性能保持。本文回顾了检索增强生成、多步推理、外部工具使用及计算表达能力等方面的最新进展,这些能力显著提升了LLM的表现。基于此,我们提出‘彩票大模型假设’:对于给定的LLM和任务,存在一个更小的‘彩票’模型,在借助多步推理和外部工具的情况下,能实现与原始模型相当的性能。据此,我们总结了彩票模型和KV缓存压缩必须具备的关键能力,而这些能力目前被主流方法所忽视。
原文摘要 · Abstract (English)
Motivated by reducing the computational and storage costs of LLMs, model compression and KV cache compression have attracted much attention from researchers. However, current methods predominantly emphasize maintaining the performance of compressed LLMs, as measured by perplexity or simple accuracy on tasks of common sense knowledge QA and basic arithmetic reasoning. In this blog, we present a brief review of recent advancements in LLMs related to retrieval-augmented generation, multi-step reasoning, external tools, and computational expressivity, all of which substantially enhance LLM performance. Then, we propose a lottery LLM hypothesis suggesting that for a given LLM and task, there exists a smaller lottery LLM capable of producing the same performance as the original LLM with the assistance of multi-step reasoning and external tools. Based on the review of current progress in LLMs, we discuss and summarize the essential capabilities that the lottery LLM and KV cache compression must possess, which are currently overlooked in existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。