构建金融文档可解释验证基准,评估大模型在长文本中的推理能力
FinDVer: Explainable Claim Verification over Long and Hybrid-Content Financial Documents
- 针对金融长文档设计三类任务:信息提取、数值推理、知识密集推理
- GPT-4o在复杂任务中仍落后于人类专家,准确率未超75%
- 提供推理错误分析,助力改进大模型在专业领域的可信验证
我们提出FinDVer,一个专门用于评估大语言模型在理解与分析长篇混合内容金融文档时可解释性论断验证能力的综合性基准。FinDVer包含2,400个专家标注样本,分为三类子集:信息抽取、数值推理和知识密集型推理,涵盖真实金融场景中的常见问题。我们在长上下文和RAG设置下评估多种LLM的表现。结果显示,即使是最先进的系统GPT-4o,其表现仍显著低于人类专家。我们进一步深入分析了长上下文与RAG设置、思维链推理以及模型推理错误,为未来提升提供了重要洞见。我们认为FinDVer可作为评估大模型在复杂专业文档中论断验证能力的重要基准。
原文摘要 · Abstract (English)
We introduce FinDVer, a comprehensive benchmark specifically designed to evaluate the explainable claim verification capabilities of LLMs in the context of understanding and analyzing long, hybrid-content financial documents. FinDVer contains 2,400 expert-annotated examples, divided into three subsets: information extraction, numerical reasoning, and knowledge-intensive reasoning, each addressing common scenarios encountered in real-world financial contexts. We assess a broad spectrum of LLMs under long-context and RAG settings. Our results show that even the current best-performing system, GPT-4o, still lags behind human experts. We further provide in-depth analysis on long-context and RAG setting, Chain-of-Thought reasoning, and model reasoning errors, offering insights to drive future advancements. We believe that FinDVer can serve as a valuable benchmark for evaluating LLMs in claim verification over complex, expert-domain documents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。