arXiv:2412.01690cs.CL2024-12中稿 · Coling 2025被引 4

提出新指标评估提示词成本与准确率平衡,发现简单方法更省钱

Can We Afford The Perfect Prompt? Balancing Cost and Accuracy with the Economical Prompting Index

  • 用综合准确率和令牌消耗的指数衡量提示成本效益
  • 复杂方法如Self-Consistency在高模型上反而更贵且提升不显著
  • 适合关注部署成本的研究者和实际应用开发者

随着提示工程研究快速发展,仅评估准确率已不足,需兼顾成本。本文提出经济提示指数(EPI),将准确率与令牌消耗结合,并通过用户设定的成本关注水平反映不同资源约束。我们在10个主流语言模型和4个多样化数据集上评估了6种先进提示技术,包括思维链(Chain-of-Thought)、自洽性(Self-Consistency)和思维树(Tree of Thoughts)。结果显示,如Self-Consistency等复杂方法常带来统计上不显著的提升,却导致成本激增。例如,在Claude 3.5 Sonnet等高性能模型上,简单方法如思维链的EPI为0.72,高于自洽性方法的0.64,即便在轻微成本关注水平下也更优。研究建议在资源受限场景下重新评估复杂提示策略,或改变未来研究方向,提升终端用户的成本效益。

原文摘要 · Abstract (English)

As prompt engineering research rapidly evolves, evaluations beyond accuracy are crucial for developing cost-effective techniques. We present the Economical Prompting Index (EPI), a novel metric that combines accuracy scores with token consumption, adjusted by a user-specified cost concern level to reflect different resource constraints. Our study examines 6 advanced prompting techniques, including Chain-of-Thought, Self-Consistency, and Tree of Thoughts, across 10 widely-used language models and 4 diverse datasets. We demonstrate that approaches such as Self-Consistency often provide statistically insignificant gains while becoming cost-prohibitive. For example, on high-performing models like Claude 3.5 Sonnet, the EPI of simpler techniques like Chain-of-Thought (0.72) surpasses more complex methods like Self-Consistency (0.64) at slight cost concern levels. Our findings suggest a reevaluation of complex prompting strategies in resource-constrained scenarios, potentially reshaping future research priorities and improving cost-effectiveness for end-users.

提示工程成本优化模型效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。