评测3大模型在金融报告生成中的表现,提供可复现的评估框架。
Evaluating Large Language Models on Financial Report Summarization: An Empirical Study
- 对比GLM-4、Mistral-NeMo、LLaMA3.1在金融报告生成中的表现
- 引入ROUGE-1、BERT Score和LLM Score等多维指标进行量化评估
- 开源金融数据集,推动领域内可复现研究与协作
近年来,大型语言模型(LLMs)在自然语言理解、领域知识任务等方面展现出卓越能力。然而,在金融等高风险高要求领域应用时,仍需严格评估其可靠性、准确性和合规性。为此,我们对三种前沿LLM——GLM-4、Mistral-NeMo和LLaMA3.1进行了全面比较研究,聚焦其在自动生成财务报告中的有效性。研究旨在探索这些模型在金融场景中的潜力,该领域对精确性、上下文相关性及抗错误信息能力要求极高。我们提出包含ROUGE-1、BERT Score和LLM Score在内的评估指标,并构建融合定量(如精度、召回率)与定性分析(如上下文契合度、一致性)的综合评估框架,全面评估各模型输出质量。此外,我们公开了金融数据集,发布于Hugging Face,欢迎研究人员和从业者使用、检验并共同改进研究成果。
原文摘要 · Abstract (English)
In recent years, Large Language Models (LLMs) have demonstrated remarkable versatility across various applications, including natural language understanding, domain-specific knowledge tasks, etc. However, applying LLMs to complex, high-stakes domains like finance requires rigorous evaluation to ensure reliability, accuracy, and compliance with industry standards. To address this need, we conduct a comprehensive and comparative study on three state-of-the-art LLMs, GLM-4, Mistral-NeMo, and LLaMA3.1, focusing on their effectiveness in generating automated financial reports. Our primary motivation is to explore how these models can be harnessed within finance, a field demanding precision, contextual relevance, and robustness against erroneous or misleading information. By examining each model's capabilities, we aim to provide an insightful assessment of their strengths and limitations. Our paper offers benchmarks for financial report analysis, encompassing proposed metrics such as ROUGE-1, BERT Score, and LLM Score. We introduce an innovative evaluation framework that integrates both quantitative metrics (e.g., precision, recall) and qualitative analyses (e.g., contextual fit, consistency) to provide a holistic view of each model's output quality. Additionally, we make our financial dataset publicly available, inviting researchers and practitioners to leverage, scrutinize, and enhance our findings through broader community engagement and collaborative improvement. Our dataset is available on huggingface.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。