压缩长提示以降低大模型推理开销,提升效率。
Prompt Compression for Large Language Models: A Survey
- 按硬提示与软提示分类,系统梳理压缩方法。
- 揭示注意力优化、参数高效微调等机制作用。
- 适合关注大模型部署效率的研究者与工程师。
利用大语言模型(LLMs)完成复杂自然语言任务通常需要长提示来表达详细要求和信息,这导致内存占用和推理成本上升。为缓解此问题,多种高效方法被提出,其中提示压缩受到广泛关注。本文综述了提示压缩技术,分为硬提示方法与软提示方法两类,对比其技术路径,并从注意力优化、参数高效微调(PEFT)、模态融合及新型合成语言等视角探讨其内在机制。同时分析各类压缩方法在下游任务中的适应性,总结当前方法的局限性,并提出未来方向,如优化压缩编码器、融合硬/软提示方法、借鉴多模态洞察。
原文摘要 · Abstract (English)
Leveraging large language models (LLMs) for complex natural language tasks typically requires long-form prompts to convey detailed requirements and information, which results in increased memory usage and inference costs. To mitigate these challenges, multiple efficient methods have been proposed, with prompt compression gaining significant research interest. This survey provides an overview of prompt compression techniques, categorized into hard prompt methods and soft prompt methods. First, the technical approaches of these methods are compared, followed by an exploration of various ways to understand their mechanisms, including the perspectives of attention optimization, Parameter-Efficient Fine-Tuning (PEFT), modality integration, and new synthetic language. We also examine the downstream adaptations of various prompt compression techniques. Finally, the limitations of current prompt compression methods are analyzed, and several future directions are outlined, such as optimizing the compression encoder, combining hard and soft prompts methods, and leveraging insights from multimodality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。