对比四款大模型在发票分类任务中的表现,找最优性价比方案。
Analysis of LLM Performance on AWS Bedrock: Receipt-item Categorisation Case Study
- 在AWS Bedrock上评估四款指令微调模型的分类性能。
- Claude 3.7 Sonnet在准确率与成本间平衡最优。
- 验证了零样本与少样本提示对效果和成本的影响。
本文针对生产环境中基于AWS Bedrock的发票项目分类任务,系统性地评估了四种指令微调大语言模型(LLMs)的表现。对比模型包括Claude 3.7 Sonnet、Claude 4 Sonnet、Mixtral 8x7B Instruct和Mistral 7B Instruct。研究目标为:(1) 评估模型在分类准确率、响应稳定性及每令牌成本方面的表现;(2) 探究零样本与少样本提示方法在准确率与成本上的适用性。实验结果表明,Claude 3.7 Sonnet在分类准确率与成本效率之间实现了最佳平衡。
原文摘要 · Abstract (English)
This paper presents a systematic, cost-aware evaluation of large language models (LLMs) for receipt-item categorisation within a production-oriented classification framework. We compare four instruction-tuned models available through AWS Bedrock: Claude 3.7 Sonnet, Claude 4 Sonnet, Mixtral 8x7B Instruct, and Mistral 7B Instruct. The aim of the study was (1) to assess performance across accuracy, response stability, and token-level cost, and (2) to investigate what prompting methods, zero-shot or few-shot, are especially appropriate both in terms of accuracy and in terms of incurred costs. Results of our experiments demonstrated that Claude 3.7 Sonnet achieves the most favourable balance between classification accuracy and cost efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。