arXiv:2510.25409cs.CLcs.AI2025-10被引 7

首个面向印度本土知识体系的多任务双语评测基准,覆盖农业、法律等四大领域。

BhashaBench V1: A Comprehensive Benchmark for the Quadrant of Indic Domains

  • 构建涵盖7.4万题目的双语问答集,分属90多个子领域
  • 发现模型在阿育吠陀等低资源领域准确率仅59.74%
  • 适合关注本地化AI评估与跨语言知识理解的研究者

大型语言模型的快速发展加剧了对特定领域与文化评估的需求。现有基准大多以英语为中心且领域无关,难以适配印度本土场景。为此,我们提出BhashaBench V1,首个聚焦印度知识体系的领域特定、多任务、双语评测基准。该基准包含74,166条精心筛选的问答对,其中52,494条为英文,21,672条为印地语,数据源自真实政府及专业考试。覆盖农业、法律、金融、阿育吠陀四大领域,涵盖90多个子领域和500多个主题,支持细粒度评估。对29+个LLMs的评估显示显著的领域与语言差异:例如,GPT-4o在法律领域整体准确率达76.49%,但在阿育吠陀领域仅59.74%;所有领域中模型在英文内容上的表现均优于印地语。子域分析表明,网络法、国际金融表现较好,而五种疗法、种子科学与人权等领域仍明显薄弱。BhashaBench V1为评估大模型在印度多元知识体系中的表现提供全面数据集,支持其整合领域知识与双语理解能力的评估。所有代码、基准与资源均公开,促进开放研究。

原文摘要 · Abstract (English)

The rapid advancement of large language models(LLMs) has intensified the need for domain and culture specific evaluation. Existing benchmarks are largely Anglocentric and domain-agnostic, limiting their applicability to India-centric contexts. To address this gap, we introduce BhashaBench V1, the first domain-specific, multi-task, bilingual benchmark focusing on critical Indic knowledge systems. BhashaBench V1 contains 74,166 meticulously curated question-answer pairs, with 52,494 in English and 21,672 in Hindi, sourced from authentic government and domain-specific exams. It spans four major domains: Agriculture, Legal, Finance, and Ayurveda, comprising 90+ subdomains and covering 500+ topics, enabling fine-grained evaluation. Evaluation of 29+ LLMs reveals significant domain and language specific performance gaps, with especially large disparities in low-resource domains. For instance, GPT-4o achieves 76.49% overall accuracy in Legal but only 59.74% in Ayurveda. Models consistently perform better on English content compared to Hindi across all domains. Subdomain-level analysis shows that areas such as Cyber Law, International Finance perform relatively well, while Panchakarma, Seed Science, and Human Rights remain notably weak. BhashaBench V1 provides a comprehensive dataset for evaluating large language models across India's diverse knowledge domains. It enables assessment of models' ability to integrate domain-specific knowledge with bilingual understanding. All code, benchmarks, and resources are publicly available to support open research.

多语言评测领域基准印度知识双语模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。