arXiv:2507.08015cs.CLcs.LG2025-07被引 1

FinGPT在金融文本分类表现强,但问答与生成任务仍逊于GPT-4。

Assessing the Capabilities and Limitations of FinGPT Model in Financial NLP Applications

  • 针对六大金融NLP任务评估,聚焦真实场景数据
  • 分类任务接近GPT-4水平,但问答与摘要性能明显不足
  • 适合用于结构化金融文本处理,不适用于复杂推理

本研究在六个关键自然语言处理任务上评估了金融领域专用语言模型FinGPT:情感分析、文本分类、命名实体识别、金融问答、文本摘要和股价走势预测。采用金融领域专用数据集,检验其在实际金融应用中的能力与局限。结果表明,FinGPT在情感分析和新闻标题分类等分类任务中表现优异,常达到与GPT-4相当的水平;但在需要推理与生成的任务如金融问答和摘要生成中表现显著下降。与GPT-4及人工基准对比显示,在数值准确性和复杂推理方面存在明显差距。总体而言,尽管FinGPT在特定结构化金融任务中有效,尚不能作为全面解决方案。该研究为未来研究提供了重要基准,并强调需对金融语言模型进行架构优化与领域适配。

原文摘要 · Abstract (English)

This work evaluates FinGPT, a financial domain-specific language model, across six key natural language processing (NLP) tasks: Sentiment Analysis, Text Classification, Named Entity Recognition, Financial Question Answering, Text Summarization, and Stock Movement Prediction. The evaluation uses finance-specific datasets to assess FinGPT's capabilities and limitations in real-world financial applications. The results show that FinGPT performs strongly in classification tasks such as sentiment analysis and headline categorization, often achieving results comparable to GPT-4. However, its performance is significantly lower in tasks that involve reasoning and generation, such as financial question answering and summarization. Comparisons with GPT-4 and human benchmarks highlight notable performance gaps, particularly in numerical accuracy and complex reasoning. Overall, the findings indicate that while FinGPT is effective for certain structured financial tasks, it is not yet a comprehensive solution. This research provides a useful benchmark for future research and underscores the need for architectural improvements and domain-specific optimization in financial language models.

金融NLP语言模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。