首个多语言多模态金融评测基准,测试大模型跨语言跨模态推理能力。
MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application
- 构建五语言、文本/视觉/音频多模态金融评测集,覆盖真实场景
- 前沿模型如GPT-4o在多语言任务中准确率仅46.01%,跨模态表现不均
- 专为专家级金融推理设计,适合评估模型在真实金融场景中的泛化能力
现实金融分析涉及多语言与多模态信息,涵盖报告、新闻、扫描文件及会议录音。然而现有金融领域大模型评测仍以单语种文本为主,且趋于饱和。为此,我们提出MultiFinBen——首个由专家标注的多语言(五种语言)和多模态(文本、视觉、音频)金融评测基准,用于评估大模型在真实金融场景下的表现。MultiFinBen引入两类新任务:多语言金融推理,检验从文件与新闻中整合跨语言证据的能力;金融OCR,从含表格和图表的扫描文档中提取结构化文本。我们采用基于模型性能与难度感知的结构化数据集筛选策略,避免冗余任务并保持挑战平衡。对21个主流大模型的评估显示,即使前沿多模态模型GPT-4o整体准确率也仅为46.01%,在视觉与音频任务中表现较强,但在多语言场景下显著下降。结果揭示了当前模型在多语言、多模态及专家级金融推理方面的持续局限。所有数据集、评估脚本与排行榜均已公开。
原文摘要 · Abstract (English)
Real-world financial analysis involves information across multiple languages and modalities, from reports and news to scanned filings and meeting recordings. Yet most existing evaluations of LLMs in finance remain text-only, monolingual, and largely saturated by current models. To bridge these gaps, we present MultiFinBen, the first expert-annotated multilingual (five languages) and multimodal (text, vision, audio) benchmark for evaluating LLMs in realistic financial contexts. MultiFinBen introduces two new task families: multilingual financial reasoning, which tests cross-lingual evidence integration from filings and news, and financial OCR, which extracts structured text from scanned documents containing tables and charts. Rather than aggregating all available datasets, we apply a structured, difficulty-aware selection based on advanced model performance, ensuring balanced challenge and removing redundant tasks. Evaluating 21 leading LLMs shows that even frontier multimodal models like GPT-4o achieve only 46.01% overall, stronger on vision and audio but dropping sharply in multilingual settings. These findings expose persistent limitations in multilingual, multimodal, and expert-level financial reasoning. All datasets, evaluation scripts, and leaderboards are publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。