构建分层评测体系,精准诊断金融大模型的分析短板。
FinReasoning: A Hierarchical Benchmark for Reliable Financial Research Reporting
- 将金融分析拆解为语义一致、数据对齐、深度洞察三层次评估
- 闭源模型整体表现最优,开源通用模型在一致性上明显不足
- 适用于多智能体系统中角色分工与模型选型决策
大语言模型正从辅助人类分析师转向多智能体自主协作,在金融研究中仍存在事实错误、数值不一致和浅层分析等问题,可能导致企业基本面误判并引发重大经济损失。现有基准对生成内容进行单一评分,无法区分模型在审计纠错等基础环节或生成研究级洞察上的具体缺陷,掩盖了能力瓶颈与专长差异。为此,我们提出FinReasoning,一个分层评测基准,将金融研究核心能力分解为语义一致性、数据对齐与深度洞察三维度,并设计细粒度评估框架,强化幻觉修正检测,引入12项指标衡量核心分析技能。评测显示,闭源模型(如Doubao-Seed-1.8)整体表现优异,适合作为多智能体系统中的核心推理代理;开源通用模型(如Qwen3-235B)在语义一致性上显著落后,不适宜质量敏感任务;金融领域模型(如Fin-R1)虽能生成中等水平洞察,但缺乏基础审计能力。该工作已在多个真实场景试点部署,资源已公开于https://github.com/TongjiFinLab/FinReasoning。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in financial research workflows, where their role is evolving from single-model assistance for human analysts toward autonomous collaboration among multiple agents. Yet real-world deployments still expose factual errors, numerical inconsistencies, and shallow analysis, which can distort assessments of corporate fundamentals and trigger severe economic losses. While existing benchmarks have begun to evaluate such failures, they score all aspects of the generated analysis in one pass, failing to distinguish whether a model fails at foundational stages like auditing and correction, or underperforms at generating research-grade insights. Consequently, it obscures capability bottlenecks and the specialized strengths essential for multi-agent role assignment. To address these gaps, we introduce FinReasoning, a hierarchical benchmark that decomposes the core capabilities of financial research into semantic consistency, data alignment, and deep insight. We further propose a fine-grained evaluation framework that strengthens hallucination-correction assessment and incorporates a 12-indicator rubric for core analytical skills. FinReasoning reveals clear capability stratification across model types. Closed-source models (like Doubao-Seed-1.8) perform strongly overall and are better suited for core reasoning agents in multi-agent financial systems; open-source general models (like Qwen3-235B) show clear capability divergence and consistently underperform on Semantic Consistency, making them less suited for quality-sensitive generation tasks; financial-domain models (like Fin-R1) generate moderate insights but lack foundational auditing skills. Our work has already been deployed in pilot tests across several real-world scenarios. The resource is available at https://github.com/TongjiFinLab/FinReasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。