用低成本预测大模型适配效果,省时省力
COSMOS: Predictable and Cost-Effective Adaptation of LLMs
- 构建统一预测框架,通过轻量代理模型预估微调性能
- 在8个基准上平均降本92.72%,极端场景最高省98.71%
- 适合资源受限下快速选模型与训练策略的研究者
大语言模型通过多种适配策略在众多任务中表现卓越,但在资源受限条件下,最优模型与策略的选择极具挑战,通常需大量实验。本文探究是否能在不进行昂贵试错的前提下,准确预测性能与成本。我们形式化了大模型适配策略选择问题,提出COSMOS统一预测框架,可高效、低成本地估计适配结果。通过两个强大预测器实现:基于嵌入增强的轻量级代理模型预测微调性能,以及低样本缩放律预测检索增强的上下文学习效果。在8个代表性基准上的广泛评估表明,COSMOS在保持高预测精度的同时,平均降低92.72%的计算成本,资源密集场景下最高可达98.71%。结果证明,高效预测适配结果不仅可行,还能显著降低大模型部署的计算开销,同时维持性能标准。
原文摘要 · Abstract (English)
Large language models (LLMs) achieve remarkable performance across numerous tasks by using a diverse array of adaptation strategies. However, optimally selecting a model and adaptation strategy under resource constraints is challenging and often requires extensive experimentation. We investigate whether it is possible to accurately predict both performance and cost without expensive trials. We formalize the strategy selection problem for LLMs and introduce COSMOS, a unified prediction framework that efficiently estimates adaptation outcomes at minimal cost. We instantiate and study the capability of our framework via a pair of powerful predictors: embedding-augmented lightweight proxy models to predict fine-tuning performance, and low-sample scaling laws to forecast retrieval-augmented in-context learning. Extensive evaluation across eight representative benchmarks demonstrates that COSMOS achieves high prediction accuracy while reducing computational costs by 92.72% on average, and up to 98.71% in resource-intensive scenarios. Our results show that efficient prediction of adaptation outcomes is not only feasible but can substantially reduce the computational overhead of LLM deployment while maintaining performance standards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。