arXiv:2602.10886cs.CLcs.AI2026-02被引 4

首个多语言多模态金融AI评测框架,覆盖理解、推理与决策全流程。

The CLEF-2026 FinMMEval Lab: Multilingual and Multimodal Evaluation of Financial AI Systems

  • 构建多语言多模态金融LLM评测体系,涵盖三类核心任务。
  • 支持跨语言、跨模态的金融理解与决策能力评估,推动全球包容性发展。
  • 面向金融AI研究者,助力开发更鲁棒、透明的智能系统。

本文介绍CLEF 2026中FinMMEval实验室的设置与任务设计,提出首个针对金融大语言模型(LLMs)的多语言、多模态评估框架。尽管近年来金融自然语言处理在市场报告、监管文件和投资者沟通的自动化分析方面取得进展,但现有基准仍以单语种、纯文本为主,且局限于特定子任务。FinMMEval 2026通过三个相互关联的任务填补这一空白:金融考试问答、多语言金融问答(PolyFiQA)以及金融决策制定。这些任务共同构成一套全面的评估体系,衡量模型在多种语言和模态下的推理、泛化与决策能力。该实验室旨在推动稳健、透明且具有全球包容性的金融AI系统发展,并公开发布数据集与评估资源,支持可复现的研究。

原文摘要 · Abstract (English)

We present the setup and the tasks of the FinMMEval Lab at CLEF 2026, which introduces the first multilingual and multimodal evaluation framework for financial Large Language Models (LLMs). While recent advances in financial natural language processing have enabled automated analysis of market reports, regulatory documents, and investor communications, existing benchmarks remain largely monolingual, text-only, and limited to narrow subtasks. FinMMEval 2026 addresses this gap by offering three interconnected tasks that span financial understanding, reasoning, and decision-making: Financial Exam Question Answering, Multilingual Financial Question Answering (PolyFiQA), and Financial Decision Making. Together, these tasks provide a comprehensive evaluation suite that measures models' ability to reason, generalize, and act across diverse languages and modalities. The lab aims to promote the development of robust, transparent, and globally inclusive financial AI systems, with datasets and evaluation resources publicly released to support reproducible research.

金融AI多语言多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。