评估提示策略时,效率比单纯性能更重要。
Incorporating Token Usage into Prompting Strategy Evaluation
- 提出Big-O_{tok}理论框架,分析提示策略的令牌消耗增长模式。
- 发现令牌使用量增加后,性能提升效果急剧下降。
- 适合关注模型推理效率与成本优化的研究者和开发者。
近年来,大语言模型在各类任务中表现出色。然而,其任务表现高度依赖于提示策略,而不同策略在性能和令牌使用量上差异显著。尽管通常以任务性能衡量提示策略优劣,但本文认为效率——即性能与令牌使用之间的平衡——更适用于实际应用。为此,我们提出Big-$O_{tok}$理论框架,用于描述提示策略的令牌使用增长特性,并引入实证指标Token Cost(每单位性能的令牌数)。对多种常见提示策略的应用分析表明,随着令牌使用量增加,性能收益呈现急剧递减趋势。结果验证了Big-$O_{tok}$分析的有效性,强调了在评估中纳入效率考量的必要性。
原文摘要 · Abstract (English)
In recent years, large language models have demonstrated remarkable performance across diverse tasks. However, their task effectiveness is heavily dependent on the prompting strategy used to elicit output, which can vary widely in both performance and token usage. While task performance is often used to determine prompting strategy success, we argue that efficiency--balancing performance and token usage--can be a more practical metric for real-world utility. To enable this, we propose Big-$O_{tok}$, a theoretical framework for describing the token usage growth of prompting strategies, and analyze Token Cost, an empirical measure of tokens per performance. We apply these to several common prompting strategies and find that increased token usage leads to drastically diminishing performance returns. Our results validate the Big-$O_{tok}$ analyses and reinforce the need for efficiency-aware evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。