用印度谜语测试多语言大模型的推理与自省能力,发现越强的模型越不自知。
The Riddle of Reflection: Evaluating Reasoning and Self-Awareness in Multilingual LLMs using Indian Riddles
- 构建多语言谜题数据集,评估七种印度语言下的模型表现。
- 顶级模型准确率超90%,但误判时自我识别率仅4.34%。
- 低性能模型反而更自知,揭示大模型自省机制缺陷。
大型语言模型(LLMs)在非英语语言中进行文化相关推理的能力仍待深入研究。本文针对印地语、泰米尔语、卡纳达语等七种主要印度语言,评估了五种大模型(Gemini 2.5 Pro、Gemini 2.5 Flash、Mistral-Saba、LLaMA 4 Scout、LLaMA 4 Maverick)在七种提示策略下的推理与自我评估能力。第一阶段评估谜题解答表现,结果显示尽管Gemini 2.5 Pro整体最优,但少样本提示仅带来微弱提升,且准确率在不同语言间差异显著。第二阶段开展自我评估实验,发现模型初始准确率与其错误识别能力呈负相关:顶尖模型(如Gemini 2.5 Pro)过于自信(真阴性率仅4.34%),而表现较差的模型(如LLaMA 4 Scout)则更具备自省能力(真阴性率达42.09%)。结果表明,当前多语言模型在推理与自我认知方面存在明显缺口,亟需发展既能有效推理又能正确认识自身局限的模型。
原文摘要 · Abstract (English)
The extent to which large language models (LLMs) can perform culturally grounded reasoning across non-English languages remains underexplored. This paper examines the reasoning and self-assessment abilities of LLMs across seven major Indian languages-Bengali, Gujarati, Hindi, Kannada, Malayalam, Tamil, and Telugu. We introduce a multilingual riddle dataset combining traditional riddles with context-reconstructed variants and evaluate five LLMs-Gemini 2.5 Pro, Gemini 2.5 Flash, Mistral-Saba, LLaMA 4 Scout, and LLaMA 4 Maverick-under seven prompting strategies. In the first stage, we assess riddle-solving performance and find that while Gemini 2.5 Pro performs best overall, few-shot methods yield only marginal gains, and accuracy varies notably across languages. In the second stage, we conduct a self-evaluation experiment to measure reasoning consistency. The results reveal a key finding: a model's initial accuracy is inversely correlated with its ability to identify its own mistakes. Top-performing models such as Gemini 2.5 Pro are overconfident (4.34% True Negative Rate), whereas lower-performing models like LLaMA 4 Scout are substantially more self-aware (42.09% True Negative Rate). These results point to clear gaps in multilingual reasoning and highlight the need for models that not only reason effectively but also recognize their own limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。