arXiv:2604.19298cs.CLcs.AI2026-04

首个评估大模型处理印度金融监管文本的基准,填补非西方金融数据空白

IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text

  • 构建406个专家标注的问答对,覆盖SEBI和RBI的192份文件,含四类任务
  • 模型零样本准确率70.4%至89.7%,数值推理任务差异最大(35.9个百分点)
  • 适合关注印度金融AI、多语言合规与大模型评测的研究者使用

我们提出IndiaFinBench,据知是首个公开可用的评估大语言模型(LLM)在印度金融监管文本上表现的基准。现有金融NLP基准均来自西方财务语料(如美国证券交易委员会文件、美国财报、英文财经新闻),缺乏对非西方监管框架的覆盖。IndiaFinBench通过从证券交易所和印度储备银行(SEBI、RBI)获取的192份文件,构建了406个专家标注的问答对,涵盖四类任务:监管解读(174项)、数值推理(92项)、矛盾检测(62项)和时间推理(78项)。标注质量经模型二次验证(矛盾检测κ=0.918;150项子集整体一致性90.7%)和三轮人工标注员间一致性研究(矛盾检测κ=0.645;整体一致率77.2%;基准覆盖率达44.3%)确认。在零样本条件下评估12个模型,准确率介于70.4%(Gemma 4 E4B)至89.7%(Gemini 2.5 Flash)之间,所有模型均显著优于非专业人类基线(69.0%)。数值推理任务最具区分度,模型间差距达35.9个百分点。10,000次自举重抽样检验揭示三个统计显著不同的性能层级。数据集、评估代码及所有模型输出已开源。

原文摘要 · Abstract (English)

We introduce IndiaFinBench, to our knowledge the first publicly available evaluation benchmark for assessing large language model (LLM) performance on Indian financial regulatory text. Existing financial NLP benchmarks draw exclusively from Western financial corpora (SEC filings, US earnings reports, English-language financial news), leaving a significant gap in coverage of non-Western regulatory frameworks. IndiaFinBench addresses this gap with 406 expert-annotated question-answer pairs drawn from 192 documents sourced from the Securities and Exchange Board of India (SEBI) and the Reserve Bank of India (RBI), spanning four task types: regulatory interpretation (174 items), numerical reasoning (92 items), contradiction detection (62 items), and temporal reasoning (78 items). Annotation quality is validated through a model-based secondary pass (kappa=0.918 on contradiction detection; 90.7% overall agreement on a 150-item subset) and a 180-item human inter-annotator agreement study across three annotation rounds (kappa=0.645 on contradiction detection; 77.2% overall agreement; 44.3% benchmark coverage). We evaluate twelve models under zero-shot conditions, with accuracy ranging from 70.4% (Gemma 4 E4B) to 89.7% (Gemini 2.5 Flash). All models substantially outperform a non-specialist human baseline of 69.0%. Numerical reasoning is the most discriminative task, with a 35.9 percentage-point spread across models. Bootstrap significance testing (10,000 resamples) reveals three statistically distinct performance tiers. The dataset, evaluation code, and all model outputs are available at https://github.com/rajveerpall/IndiaFinBench

金融AI多语言评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。