arXiv:2508.19831cs.CLcs.LG2025-08被引 1

为印地语大模型打造首个高质量评测套件,解决文化语义缺失问题。

Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis

  • 自建五套印地语评测数据集,结合人工标注与翻译验证。
  • 首次系统评估开源印地语大模型性能,揭示现有模型短板。
  • 方法可复用于其他低资源语言,推动多语言AI发展。

由于缺乏高质量基准测试,评估印地语指令微调大语言模型(LLMs)面临挑战,直接翻译英文数据集无法捕捉关键的语言和文化细节。为此,我们推出五套印地语LLM评估数据集:IFEval-Hi、MT-Bench-Hi、GSM8K-Hi、ChatRAG-Hi 和 BFCL-Hi。这些数据集采用从头人工标注与翻译-验证相结合的方法构建。利用该套数据集,我们对支持印地语的开源大模型进行了全面基准测试,并提供了详细的性能对比分析。我们的数据构建流程也可作为其他低资源语言基准开发的可复现范式。

原文摘要 · Abstract (English)

Evaluating instruction-tuned Large Language Models (LLMs) in Hindi is challenging due to a lack of high-quality benchmarks, as direct translation of English datasets fails to capture crucial linguistic and cultural nuances. To address this, we introduce a suite of five Hindi LLM evaluation datasets: IFEval-Hi, MT-Bench-Hi, GSM8K-Hi, ChatRAG-Hi, and BFCL-Hi. These were created using a methodology that combines from-scratch human annotation with a translate-and-verify process. We leverage this suite to conduct an extensive benchmarking of open-source LLMs supporting Hindi, providing a detailed comparative analysis of their current capabilities. Our curation process also serves as a replicable methodology for developing benchmarks in other low-resource languages.

印地语大模型评测低资源语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。