arXiv:2507.22936cs.CLcs.AI2025-07被引 6

对比五种大模型在财报问答中的表现,发现无一模型全面领先。

Evaluating Large Language Models (LLMs) in Financial NLP: A Comparative Study on Financial Report Analysis

  • 采用人类评估、自动相似度与行为诊断三重方法综合评测
  • 各模型在相关性、准确性等维度表现差异明显,无统一最优
  • 提醒金融场景部署需考虑主观判断与模型行为波动性

大语言模型在复杂财务披露分析中的应用日益广泛,但其可靠性、行为一致性与透明度在高风险场景中仍不充分。本文对五种基于Transformer的LLM在美国内部10-K文件业务章节上的问答任务进行了受控评估。为捕捉模型行为的多维度特征,结合了人工评估、自动化相似度指标及标准化提示下的行为诊断。人工评估显示,模型在相关性、完整性、清晰度、简洁性和事实准确性等定性维度上表现各异,但评估者间一致性较低,反映标准主观性。自动指标揭示模型间在词汇重叠和语义相似度上存在系统性差异,行为诊断则指出响应稳定性与跨提示一致性存在显著变异。重要的是,无单一模型在所有评估视角中持续领先。结果表明,性能差异应理解为特定条件下的相对倾向,而非普遍可靠性指标。研究强调,在财务关键应用中部署LLM时,需构建能处理人类分歧、行为变异性与可解释性的评估框架。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to support the analysis of complex financial disclosures, yet their reliability, behavioral consistency, and transparency remain insufficiently understood in high-stakes settings. This paper presents a controlled evaluation of five transformer-based LLMs applied to question answering over the Business sections of U.S. 10-K filings. To capture complementary aspects of model behavior, we combine human evaluation, automated similarity metrics, and behavioral diagnostics under standardized and context-controlled prompting conditions. Human assessments indicate that models differ in their average performance across qualitative dimensions such as relevance, completeness, clarity, conciseness, and factual accuracy, though inter-rater agreement is modest, reflecting the subjective nature of these criteria. Automated metrics reveal systematic differences in lexical overlap and semantic similarity across models, while behavioral diagnostics highlight variation in response stability and cross-prompt alignment. Importantly, no single model consistently dominates across all evaluation perspectives. Together, these findings suggest that apparent performance differences should be interpreted as relative tendencies under the tested conditions rather than definitive indicators of general reliability. The results underscore the need for evaluation frameworks that account for human disagreement, behavioral variability, and interpretability when deploying LLMs in financially consequential applications.

金融NLP大模型评估10-K报告多维评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。