arXiv:2601.22162q-fin.GNcs.AI2026-01被引 1

首个面向金融多模态模型的统一评估基准,覆盖文本、图像、视频。

UniFinEval: Towards Unified Evaluation of Financial Multimodal Models across Text, Images and Videos

  • 构建涵盖五类真实金融场景的多模态评测集
  • 10个主流模型在零样本与思维链设置下表现均未达专家水平
  • 适合金融AI研究者与多模态模型开发者参考

多模态大语言模型在金融领域日益重要,但其面临的信息密度高、跨模态多跳推理等挑战超出了现有基准的评估范围。为此,我们提出UniFinEval,首个专为高信息密度金融环境设计的统一多模态评测基准,覆盖文本、图像和视频。该基准系统性构建了五个基于真实金融系统的核心场景:财务报表审计、公司基本面推理、行业趋势洞察、金融风险感知与资产配置分析。我们人工构建了一个包含3,767对问答数据的高质量数据集,支持中英文,对10个主流多模态大模型在零样本与思维链(CoT)设置下进行系统评估。结果表明,Gemini-3-pro-preview表现最佳,但仍与金融专家存在显著差距。进一步错误分析揭示当前模型存在系统性缺陷。UniFinEval旨在全面评估多模态大模型在细粒度、高信息密度金融环境中的能力,提升其在真实金融场景应用的鲁棒性。数据与代码已开源:https://github.com/aifinlab/UniFinEval。

原文摘要 · Abstract (English)

Multimodal large language models are playing an increasingly significant role in empowering the financial domain, however, the challenges they face, such as multimodal and high-density information and cross-modal multi-hop reasoning, go beyond the evaluation scope of existing multimodal benchmarks. To address this gap, we propose UniFinEval, the first unified multimodal benchmark designed for high-information-density financial environments, covering text, images, and videos. UniFinEval systematically constructs five core financial scenarios grounded in real-world financial systems: Financial Statement Auditing, Company Fundamental Reasoning, Industry Trend Insights, Financial Risk Sensing, and Asset Allocation Analysis. We manually construct a high-quality dataset consisting of 3,767 question-answer pairs in both chinese and english and systematically evaluate 10 mainstream MLLMs under Zero-Shot and CoT settings. Results show that Gemini-3-pro-preview achieves the best overall performance, yet still exhibits a substantial gap compared to financial experts. Further error analysis reveals systematic deficiencies in current models. UniFinEval aims to provide a systematic assessment of MLLMs' capabilities in fine-grained, high-information-density financial environments, thereby enhancing the robustness of MLLMs applications in real-world financial scenarios. Data and code are available at https://github.com/aifinlab/UniFinEval.

多模态金融AI评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。