测试印度语与混语间知识一致性,发现混语能显著提升模型表现。
Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR

- 构建多语言混合数据集IndicKLAR,覆盖18种印度语及11组混语对。
- 混语输入使模型性能接近英语水平,误差仅0.05左右。
- 内部语言转换策略比外部翻译更有效,且存在稳定预测转折点。
大型语言模型在英语中知识召回可靠,但在低资源语言中常失败——这一跨语言一致性差距在印度语及其混语形式中仍缺乏研究。为此,我们提出IndicKLAR,作为KLAR-CLC基准的印地语扩展,涵盖22种官方印度语言中的18种,并为11组常用语言对构建混语变体,所有设置均经母语者验证。该三向对齐为分析英语、混语与原生印度语输入下的知识召回一致性提供了独特机会。在九个开源模型上评估发现,原生语言准确率与英语差距可达约0.50,而混语输入可大幅缩小此差距,使性能逼近英语水平(误差约0.05),无需模型级修改。进一步测试多种提示策略:两阶段翻译后回答、单阶段联合翻译与回答,以及TinT——一种模型内部完成语言转换并仅输出最终答案的单步策略。在从原生语言→混语→英语的性能轨迹中,识别出一个一致的翻转点,即正确与错误预测的分界线,该点位于原生与混语之间。有趣的是,无论通过输入表面形式还是模型内部转换过程诱导,该翻转点均成立。
原文摘要 · Abstract (English)
Large language models recall knowledge reliably in English but often fail on the same query posed in a lower-resourced language -- a crosslingual consistency gap that remains underexplored for Indian languages and their code-mixed counterparts. To study this gap, we introduce IndiKLAR, an Indic extension of the KLAR-CLC benchmark covering 18 of the 22 scheduled Indian languages and pairing them with code-mixed variants for 11 widely used language pairs, with native-speaker verification of both monolingual and code-mixed variants for these 11 settings. This three-way alignment offers a unique opportunity to examine how knowledge recall consistency varies across the spectrum of English, code-mixed, and native Indian language inputs. Evaluating across nine open-weight models, we find that the native-language accuracy gap to English can reach $\sim$0.50, while code-mixed inputs close most of it -- bringing performance within $\sim$0.05 of English without any model-level intervention. Motivated by this, we evaluate several prompting strategies that vary in how language conversion is exposed, including a two-stage translate-then-answer setup, a one-stage joint translation-and-answer prompt, and Translate-in-Thought (TinT) -- a single-step strategy in which the model converts the input internally and emits only the final answer. Across the performance trajectory native $\rightarrow$ code-mixed $\rightarrow$ English, we identify a consistent flip point -- the boundary between incorrect and correct prediction -- that lies between the native and code-mixed settings. Interestingly, this holds whether the trajectory is induced by the input surface form or by the model's internal conversion process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。