首个面向金融信贷的多模态基准,专为真实场景设计
FCMBench: The First Large-scale Financial Credit Multimodal Benchmark for Real-world Applications
- 构建覆盖26类证书的隐私合规数据集,含5198张图像和13806对VQA样本
- 评估模型在感知与推理任务中面对10类真实干扰时的鲁棒性表现
- 适合金融、AI安全、多模态系统研究者使用,推动真实场景应用
FCMBench 是首个面向真实金融信贷场景的大规模、隐私合规多模态基准,涵盖领域特定工作流与约束下的各类任务与鲁棒性挑战。当前版本包含26种证件类型,共5198张隐私合规图像和13806对配对的VQA样本。该基准评估模型在3项基础感知任务、4项需决策导向视觉证据解析的信贷特定推理任务,以及10项真实世界干扰下的严格鲁棒性压力测试中的表现。数据通过自研场景感知模板人工合成,无公开图像泄露风险。我们对来自14家机构的28个先进视觉语言模型进行了广泛评估,其中商业模型中Gemini 3 Pro表现最佳(F1=65.16),开源模型中Kimi-K2.5表现最佳(F1=60.58)。所有模型平均得分44.8,标准差10.3,表明该基准具有足够难度,能有效区分模型能力。鲁棒性测试显示,即使顶尖模型在干扰下也出现显著性能下降。本基准已开源,旨在推动信贷领域AI研究与真实应用发展。
原文摘要 · Abstract (English)
FCMBench is the first large-scale and privacy-compliant multimodal benchmark for real-world financial credit applications, covering tasks and robustness challenges from domain specific workflows and constraints. The current version of FCMBench covers 26 certificate types, with 5198 privacy-compliant images and 13806 paired VQA samples. It evaluates models on Perception and Reasoning tasks under real-world Robustness interferences, including 3 foundational perception tasks, 4 credit-specific reasoning tasks demanding decision-oriented visual evidence interpretation, and 10 real-world challenges for rigorous robustness stress testing. Moreover, FCMBench offers privacy-compliant realism with minimal leakage risk through in-house scenario-aware captures of manually synthesized templates, without any publicly released images. We conduct extensive evaluations of 28 state-of-the-art vision-language models spanning 14 AI companies and research institutes. Among them, Gemini 3 Pro achieves the best F1 score as a commercial model (65.16), Kimi-K2.5 achieves the best score as an open-source baseline (60.58). The mean and the std. of all tested models is 44.8 and 10.3 respectively, indicating that FCMBench is non-trivial and provides strong resolution for separating modern vision-language model capabilities. Robustness evaluations reveal that even top-performing models experience notable performance degradation under the designed challenges. We have open-sourced this benchmark to advance AI research in the credit domain and provide a domain-specific task for real-world AI applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。