针对多语言大模型,提出动态选择提示策略的自适应方法。
No One Fits All: From Fixed Prompting to Learned Routing in Multilingual LLMs

- 将提示策略选择建模为可学习的决策问题,用轻量分类器判断每例该用原生或翻译提示。
- 在四个基准上显著优于固定策略,且能泛化到未见过的任务格式。
- 发现语言资源丰富度决定翻译是否有效,而非翻译质量本身。
多语言大模型广泛采用基于翻译的提示,但其效果在不同语言和任务间差异显著。我们在十种资源水平不同的语言及四个基准上评估了提示策略。分析表明,无单一策略始终最优:低资源语言即使翻译质量不佳也受益于翻译提示,高资源语言收益甚微,基于提示的自路由表现不如显式翻译。受此启发,我们将提示策略选择建模为可学习的决策问题,引入轻量级分类器,预测每个实例使用原生或翻译提示的最优性。该方法在四个基准上均实现统计显著提升,并能泛化至训练中未见的任务格式。进一步分析显示,语言资源水平而非翻译质量是决定翻译是否有效的关键因素。
原文摘要 · Abstract (English)
Translation-based prompting is widely used in multilingual LLMs, yet its effectiveness varies across languages and tasks. We evaluate prompting strategies across ten languages of different resource levels and four benchmarks. Our analysis shows that no single strategy is universally optimal. Translation strongly benefits low-resource languages even when translation quality is imperfect, high-resource languages gain little, and prompt-based self-routing underperforms explicit translation. Motivated by these findings, we formulate prompting strategy selection as a learned decision problem and introduce lightweight classifiers that predict whether native or translation-based prompting is optimal for each instance. The classifiers achieve statistically significant improvements over fixed strategies across four benchmarks and generalize to unseen task formats not observed during training. Further analysis reveals that language resource level, rather than translation quality alone, determines when translation is beneficial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。