对比翻译时机对跨语言提示的影响,提升低资源语言性能。
How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting
- 比较提前翻译和直接跨语言提示两种策略的优劣。
- 优化提示策略使下游分类任务性能显著提升。
- 适合关注多语言模型与低资源语言应用的研究者。
尽管大型语言模型(LLMs)在多语言能力上取得进展,其在不同语言和任务上的表现仍存在显著差异。在基于多语言检索增强生成(RAG)的系统中,知识库(KB)常从高资源语言(如英语)共享至低资源语言,导致检索信息与上下文语言不一致。此时常见的做法是预翻译以构建单语提示,或直接采用跨语言提示进行推理。然而,这些选择的影响尚不明确。本文系统评估了多语言系统中分类任务下,RAG增强型LLM所采用的不同提示翻译策略的影响。实验结果表明,优化的提示策略能显著提升跨语言知识共享效果,从而改善下游分类任务表现。研究倡导更广泛地利用多语言资源共享及跨语言提示优化,尤其针对非英语语言,特别是低资源语言。
原文摘要 · Abstract (English)
Despite advances in the multilingual capabilities of Large Language Models (LLMs), their performance varies substantially across different languages and tasks. In multilingual retrieval-augmented generation (RAG)-based systems, knowledge bases (KB) are often shared from high-resource languages (such as English) to low-resource ones, resulting in retrieved information from the KB being in a different language than the rest of the context. In such scenarios, two common practices are pre-translation to create a mono-lingual prompt and cross-lingual prompting for direct inference. However, the impact of these choices remains unclear. In this paper, we systematically evaluate the impact of different prompt translation strategies for classification tasks with RAG-enhanced LLMs in multilingual systems. Experimental results show that an optimized prompting strategy can significantly improve knowledge sharing across languages, therefore improve the performance on the downstream classification task. The findings advocate for a broader utilization of multilingual resource sharing and cross-lingual prompt optimization for non-English languages, especially the low-resource ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。