arXiv:2607.25375cs.CL2026-07

为印度语言文化量身打造的LLM评测框架,填补本土化评估空白。

Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context

  • 基于英国AI安全研究所平台构建,覆盖16种印度语言的多任务评测
  • Sarvam-M 24B和Gemma 2 27B在公平性指标上达80%,超越32B模型
  • 专设数字公共基础设施安全与文化知识测试,适合本土化模型优化

印度拥有超过14亿人口,涵盖22种官方语言和数百种地方文化传统。当前主流大语言模型评测基准(如MMLU、BIG-Bench、TruthfulQA)几乎全为英语和西方中心,无法识别印度特有的安全、公平与准确性问题。本文提出Inspect India Evals——一个开源框架,基于UK AISI的Inspect AI平台构建,包含六项评测:16种印度语言的多语言MMLU、针对印度社会偏见的BharatBBQ、数字公共基础设施(DPI)安全评估、使用印度语有害提示的多语言安全测试、多轮越狱抵抗测试,以及基于LLM-as-judge的印度文化知识基准。研究测试了5个开权重模型(8B至32B参数)。结果显示,Sarvam-M 24B和Gemma 2 27B在综合印度公平性指数上均达80%,其中Sarvam-M在印度文化知识和DPI安全合规性上超越更大模型。所有模型在多语言安全测试中均100%拒绝,但DPI安全表现从20%至100%不等。该框架已公开,支持与UK AISI注册表集成,可自由复现与扩展。

原文摘要 · Abstract (English)

India is a vast nation of over 1.4 billion people, varied by hundreds of diverse and locally specific traditions and cultures and 22 officially recognized languages. Large language models (LLMs) are now being deployed on a massive scale throughout the mainland as well as in remote villages. However, the common benchmarks - MMLU, BIG-Bench, and TruthfulQA are almost exclusively English- and Western-centric. They do not identify those safety, fairness, and accuracy failures unique to the Indian context. That is the gap Inspect India Evals seeks to fill. It is an open-source framework built on top of UK AISI's Inspect AI platform. It has six benchmarks: Multilingual MMLU across sixteen Indian languages, BharatBBQ (our adaptation of BBQ for Indian social bias), a safety evaluation for Digital Public Infrastructure, a multilingual safety test using harmful prompts in Indian languages, a multi-turn jailbreak resistance test, and an Indian cultural knowledge benchmark scored using LLM-as-judge rubrics. In this study, we tested five open-weight models ranging from 8B to 32B parameters. Sarvam-M 24B and Gemma 2 27B came out on top, both scoring 80% on the composite India Fairness Index, with Sarvam-M even beating larger 32B models on Indian cultural knowledge and DPI safety compliance. All models scored 100% refusal on Multilingual Safety, whereas DPI safety varied from 20% to 100%. The framework is public. It's built to work with the UK AISI registry. Anyone can reproduce or extend this work.

大模型评测多语言文化适配AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。