arXiv:2503.16575cs.CLcs.AI2025-03被引 4

针对长篇金融问答评估难题,提出三步法新框架。

Extract, Match, and Score: An Evaluation Paradigm for Long Question-context-answer Triplets in Financial Analysis

  • 分三步提取、匹配、打分,专为长文本设计
  • 在真实金融数据上验证,传统方法效果差
  • 适合评估复杂场景下大模型的输出质量

大语言模型(LLMs)的快速发展推动其在众多领域广泛应用,但现有评估框架在长文本场景下表现不佳。尤其在金融分析等需要处理长问题、长上下文和长回答的实际应用中,传统评价指标的适用性显著下降。本文以真实金融场景为例,构建了一个包含长篇问答三元组的数据集,并揭示了传统方法的局限性。为此,提出一种专为长文本设计的「提取-匹配-评分」(Extract, Match, and Score, EMS)评估方法,为从业者提供了一套可靠的评估工具,有效衡量大模型在复杂现实任务中的表现。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has sparked widespread adoption across diverse applications, making robust evaluation frameworks crucial for assessing their performance. While conventional evaluation metrics remain applicable for shorter texts, their efficacy diminishes when evaluating the quality of long-form answers. This limitation is particularly critical in real-world scenarios involving extended questions, extensive context, and long-form answers, such as financial analysis or regulatory compliance. In this paper, we use a practical financial use case to illustrate applications that handle "long question-context-answer triplets". We construct a real-world financial dataset comprising long triplets and demonstrate the inadequacies of traditional metrics. To address this, we propose an effective Extract, Match, and Score (EMS) evaluation approach tailored to the complexities of long-form LLMs' outputs, providing practitioners with a reliable methodology for assessing LLMs' performance in complex real-world scenarios.

金融分析大模型评估长文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。