arXiv:2605.13330cs.CL2026-05ACL被引 1

构建首个印地语金融多模态推理数据集与模型,支持多语言财务决策。

FIND: Toward Multimodal Financial Reasoning and Question Answering for Indic Languages

论文配图:FIND: Toward Multimodal Financial Reasoning and Question Answering for Indic Languages
图 1 · 摘自论文原文
  • 设计跨6种印地语系语言的多模态金融问答数据集
  • 涵盖14个金融领域共1.89万条样本,分三难度四题型
  • 提出FIND框架实现数值准确、多模态对齐的财务推理

多语言环境下的金融决策需要基于多种模态的精确数值推理,但现有基准普遍忽视这一高风险现实挑战,尤其针对印地语系语言。我们提出FinVQA,一个用于评估多语言印地语环境下金融数值与多模态推理的基准。FinVQA覆盖英语、印地语、孟加拉语、马拉地语、古吉拉特语和泰米尔语,包含18,900个样本,涵盖14个金融领域。数据集在真实约束下捕捉多样化的推理范式,分为三个难度级别(简单、中等、困难)和四种问题格式:单选题、填空题、表格匹配题和判断题。为应对这些挑战,我们提出FIND框架,结合监督微调与约束感知解码,促进忠实的数值推理、鲁棒的多模态对齐及结构化决策。FinVQA与FIND共同建立了一个严谨的高风险多语言多模态金融推理评估与建模范式。

原文摘要 · Abstract (English)

Financial decision-making in multilingual settings demands accurate numerical reasoning grounded in diverse modalities, yet existing benchmarks largely overlook this high-stakes, real-world challenge, especially for Indic languages. We introduce FinVQA, a benchmark for evaluating financial numerical and multimodal reasoning in multilingual Indic contexts. FinVQA spans English, Hindi, Bengali, Marathi, Gujarati, and Tamil, and comprises 18,900 samples across 14 financial domains. The dataset captures diverse reasoning paradigms under realistic constraints, and is structured across three difficulty levels (easy, moderate, hard) and four question formats: multiple choice, fill-in-the-blank, table matching, and true/false. To address these challenges, we propose FIND, a framework that combines supervised fine-tuning with constraint-aware decoding to promote faithful numerical reasoning, robust multimodal grounding, and structured decision-making. Together, FinVQA and FIND establish a rigorous evaluation and modeling paradigm for high-stakes multilingual multimodal financial reasoning.

金融推理多模态印地语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。