大模型跨文字体系知识迁移差,主要因文字类型不匹配。
Large Reasoning Models Struggle to Transfer Parametric Knowledge Across Scripts
- 用数据证明文字体系比语言族类更影响知识迁移
- 给模型提供源语言关键实体,显著提升跨文字问答准确率
- 通过微调训练让模型推理时更好处理音译歧义,缩小迁移差距
本文分析大型现代推理模型在跨语言知识迁移中的缺陷。通过对包含全球本地知识的ECLeKTic和MultiLoKo两个数据集进行观测分析,回归结果表明:在控制模型能力与问题难度后,文字体系匹配度是知识迁移失败的主要预测因子,而非语言或语系。进一步实验发现,向模型提供问题中关键实体的源语言形式,能显著改善跨文字问题的表现。为此,我们构建了一个合成生成管道,设计SFT样本以引导模型在推理时更好处理音译歧义,从而获取参数化知识。实验证明,对两个模型进行此类微调可有效缩小跨文字迁移差距。结论表明,后训练阶段存在提升跨语言参数知识迁移的潜力。
原文摘要 · Abstract (English)
In this work, we analyze shortcomings in cross-lingual knowledge transfer in large, modern reasoning LLMs. We demonstrate that the perceived gap in knowledge transfer is primarily a script barrier. First, we conduct an observational data analysis on the performance of thinking models on two datasets with local knowledge from around the world, ECLeKTic and MultiLoKo. Our regression analysis shows that script match - not language or family - is the primary predictor of knowledge transfer failure once model capability and question difficulty are accounted for. We further this finding by providing the LLMs with the key entities of the questions in their source language and find that this disproportionately improves cross-script questions. We then posit that these LLMs could be reasoning better at test-time. To evaluate this, we develop a synthetic generation pipeline to design SFT samples to encourage the model to better reason about transliteration ambiguities when trying to fetch parametric knowledge at inference-time. We show that teaching two models to reason better reduces the cross-script transfer gap. As a result, we conclude that there is potential to improve cross-lingual parametric knowledge transfer during post-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。