小模型也能搞定金融任务,效率还更高
Is GPT-OSS All You Need? Benchmarking Large Language Models for Financial Intelligence and the Surprising Efficiency Paradox
- 用更小的GPT-OSS-20B模型完成金融文本任务
- 准确率65.1%接近大模型,速度达159.8 tokens/秒
- 适合追求高效低成本部署的金融场景
大型语言模型在金融领域的快速应用亟需严谨的评估框架以衡量其性能、效率与实际可用性。本文对GPT-OSS系列模型及当前主流LLM在十项多样化的金融NLP任务中进行了全面评估。通过对1200亿和200亿参数版本的GPT-OSS进行大规模实验,我们发现反直觉现象:较小的GPT-OSS-20B模型在准确率(65.1%)上仅略低于大模型(66.5%),但计算效率更优,获得198.4的词元效率得分与159.80 tokens/秒的处理速度。评估涵盖情感分析、问答与实体识别,使用Financial PhraseBank、FiQA-SA、FLARE FINERORD等真实金融数据集。本文引入新型效率指标,量化性能与资源消耗的权衡,为生产环境部署提供关键参考。结果表明,GPT-OSS模型持续优于更大规模对手(如Qwen3-235B),挑战了模型规模越大性能越高的普遍认知。研究证明,GPT-OSS的架构创新与训练策略使小型模型在显著降低计算开销的前提下仍能实现竞争力表现,为金融领域可持续、低成本的LLM部署提供了可行路径。
原文摘要 · Abstract (English)
The rapid adoption of large language models in financial services necessitates rigorous evaluation frameworks to assess their performance, efficiency, and practical applicability. This paper conducts a comprehensive evaluation of the GPT-OSS model family alongside contemporary LLMs across ten diverse financial NLP tasks. Through extensive experimentation on 120B and 20B parameter variants of GPT-OSS, we reveal a counterintuitive finding: the smaller GPT-OSS-20B model achieves comparable accuracy (65.1% vs 66.5%) while demonstrating superior computational efficiency with 198.4 Token Efficiency Score and 159.80 tokens per second processing speed [1]. Our evaluation encompasses sentiment analysis, question answering, and entity recognition tasks using real-world financial datasets including Financial PhraseBank, FiQA-SA, and FLARE FINERORD. We introduce novel efficiency metrics that capture the trade-off between model performance and resource utilization, providing critical insights for deployment decisions in production environments. The benchmark reveals that GPT-OSS models consistently outperform larger competitors including Qwen3-235B, challenging the prevailing assumption that model scale directly correlates with task performance [2]. Our findings demonstrate that architectural innovations and training strategies in GPT-OSS enable smaller models to achieve competitive performance with significantly reduced computational overhead, offering a pathway toward sustainable and cost-effective deployment of LLMs in financial applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。