小模型在多数任务上比大模型更高效,适合资源受限场景。
Task-Specific Efficiency Analysis: When Small Language Models Outperform Large Language Models
- 提出性能-效率比(PER)综合评估准确率、吞吐、内存和延迟。
- 0.5–3B参数的小模型在5个任务中均取得更高PER得分。
- 为追求推理效率的生产部署提供量化依据,适合边缘计算等场景。
大语言模型虽表现优异,但计算成本高,不适用于资源受限环境。本文首次系统性地对比了16个语言模型在五个多样化自然语言处理任务上的表现。提出性能-效率比(PER)这一新指标,通过几何平均归一化整合准确率、吞吐量、内存占用与延迟。系统评估表明,0.5–3B参数的小模型在所有任务中均获得更高的PER分数。研究结果为优先考虑推理效率而非微小准确率提升的生产环境部署提供了量化基础。
原文摘要 · Abstract (English)
Large Language Models achieve remarkable performance but incur substantial computational costs unsuitable for resource-constrained deployments. This paper presents the first comprehensive task-specific efficiency analysis comparing 16 language models across five diverse NLP tasks. We introduce the Performance-Efficiency Ratio (PER), a novel metric integrating accuracy, throughput, memory, and latency through geometric mean normalization. Our systematic evaluation reveals that small models (0.5--3B parameters) achieve superior PER scores across all given tasks. These findings establish quantitative foundations for deploying small models in production environments prioritizing inference efficiency over marginal accuracy gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。