评测大模型在金融长文本问答中的可信溯源能力
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
- 设计新基准,评估模型生成答案时的证据提取与推理过程
- 8个模型测试显示,端到端生成性能不输事后修正方法
- 适合关注金融AI可信度与可解释性的研究者和从业者
大型语言模型在回答长文本问题时常出现幻觉,生成看似合理却事实错误的答案。现有评测多聚焦于简单的引用检索,但在金融等真实场景中,溯源应包含更复杂的要素。我们提出FinLFQA,一个用于评估大模型在复杂金融问题上生成长文本答案及其可靠、细致溯源能力的新基准。该基准通过人工标注评估三个关键维度:(1)从财务报告中提取的支持性证据,(2)中间的数值推理步骤,(3)支撑推理的领域特定金融知识。我们还构建了自动评估框架,涵盖答案质量与溯源质量。在八个大模型上进行多范式实验发现,细粒度指标对区分模型能力至关重要;端到端生成表现与事后修正相当;迭代优化仅在外部反馈引导下有效。
原文摘要 · Abstract (English)
Large Language Models (LLMs) frequently hallucinate to long-form questions, producing plausible yet factually incorrect answers. A common mitigation strategy is to provide attribution to LLM outputs. However, existing benchmarks primarily focus on simple attribution that retrieves supporting textual evidence as references. We argue that in real-world scenarios such as financial applications, attribution goes beyond reference retrieval. We introduce FinLFQA, a benchmark designed to evaluate the ability of LLMs to generate long-form answers to complex financial questions with reliable and nuanced attributions. FinLFQA evaluates three critical aspects of attribution through human annotations: (1) supporting evidence extracted from financial reports, (2) intermediate numerical reasoning steps, and (3) domain-specific financial knowledge that informs the reasoning process. We further provide an automatic evaluation framework covering both answer quality and attribution quality. Through extensive experiments on eight LLMs across multiple attribution-generation paradigms, we find that fine-grained metrics are important to distinguish model capabilities, that end-to-end generation achieves comparable performance to post-hoc approaches, and that iterative refinement only helps when guided by external feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。