构建首个乌尔都语推理评估基准,解决低资源语言推理评测难题
UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop
- 通过上下文融合翻译+人工验证,确保翻译后推理任务的语义一致性
- 在4个数据集、5种难度下测试模型,发现多步推理在乌尔都语中挑战巨大
- 适用于低资源语言推理评估,为多语言模型研究提供可复用框架
大语言模型虽具备强大推理能力,但低资源语言的评测仍面临标准基准缺失的挑战。尤其对乌尔都语而言,机器翻译敏感性高,且缺乏聚焦推理的任务基准。本文提出一种结合上下文融合翻译与人工验证的框架,将MGSM、MATH-500、CommonSenseQA和OpenBookQA等主流推理与问答基准翻译成乌尔都语,形成统一的UrduBench基准。我们评估了多种推理导向及指令微调的LLM,在不同提示策略下的表现。分析显示:(1)四个数据集间性能差异显著;(2)五类任务难度下模型表现不一;(3)不同模型架构与缩放设置影响明显;(4)语言一致性测试中表现不稳定。结果表明,多步与符号推理在乌尔都语中尤为困难,语言对齐稳定性是可靠推理的前提。本研究建立了一套可扩展的乌尔都语推理评测方法,并揭示了多语言推理失败的实证规律,该实验设计亦适用于其他低资源语言。代码与数据集将公开发布。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have led to strong reasoning capabilities; however, evaluating such models in low-resource languages remains challenging due to the lack of standardized benchmarks. In particular, Urdu reasoning evaluation has been limited by the sensitivity of machine translation and an emphasis on general language tasks rather than reasoning benchmarks. In this paper, we propose a contextually ensembled translation framework with human-in-the-loop validation that leverages multiple translation systems to develop Urdu reasoning benchmarks while preserving contextual and structural integrity. Using this framework, we translate widely adopted reasoning and question-answering benchmarks, including MGSM, MATH-500, CommonSenseQA, and OpenBookQA, into Urdu, collectively referred to as UrduBench, and conduct a comprehensive evaluation of both reasoning-oriented and instruction-tuned LLMs across multiple prompting strategies. Our analysis reveals performance differences across (1) four datasets, (2) five task difficulty levels, (3) diverse model architectures, (4) multiple model scaling settings, and (5) language consistency tests. We find that multi-step and symbolic reasoning tasks pose significant challenges in Urdu, and that stable language alignment is a critical prerequisite for robust reasoning. Overall, our work establishes a scalable methodology for standardized reasoning evaluation in Urdu and provides empirical insights into multilingual reasoning failures. This experimental setup is also broadly applicable to other low-resource languages. The code and datasets will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。