测试发现多语言大模型在印地语推理中表现不佳,英语提示反而更有效。
Multilingual LLMs Are Not Multilingual Thinkers: Evidence from Hindi Analogy Evaluation
- 构建405道印地语类比题数据集HATS,来自印度政府考试。
- 模型在英语提示下表现最佳,印地语提示效果差。
- 提出基于认知理论的链式思维方法,提升印地语类比推理能力。
类比测试能评估模型对概念间隐含关系的推断能力,是检验推理能力的关键基准。尽管大语言模型(LLMs)在英语推理方面被广泛研究,其在印地语等印欧语系语言中的表现仍缺乏系统评估,限制了我们对模型跨语言泛化能力的理解。为填补这一空白,我们引入新的印地语类比测试集(HATS),包含405道来自印度政府考试的多选题。我们采用多种提示策略,对主流多语言大模型进行评测,并提出一种基于认知理论的“具身链式思维”方法,显著提升模型在印地语类比题上的表现。实验表明,无论采用何种提示策略,使用英语提示时模型性能始终最优。本测试集弥补了印地语推理能力评估资源的缺失。
原文摘要 · Abstract (English)
Analogies test a model's ability to infer implicit relationships between concepts, making them a key benchmark for evaluating reasoning capabilities. While large language models (LLMs) are widely evaluated for reasoning in English, their abilities in Indic languages remain understudied, limiting our understanding of whether these models generalize across languages. To address this gap, we introduce a new Hindi Analogy Test Set (HATS), comprising 405 multiple-choice questions sourced from Indian government exams. We benchmark state-of-the-art multilingual LLMs using various prompting strategies and introduce a grounded Chain of Thought approach that leverages cognitive theories of analogical reasoning. This approach improves model performance on Hindi analogy questions. Our experiments show that models perform best with English prompts, irrespective of the prompting strategy. Our test set addresses the lack of a critical resource to evaluate LLM reasoning capabilities in Hindi.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。