优化提示词提升多语言大模型性能,效果优于单纯翻译。
The Art of Asking: Multilingual Prompt Optimization for Synthetic Data
- 通过自然度、文化适配和难度增强改造提示词,而非仅翻译。
- 在12种语言上实现全球MMLU准确率提升4.7%。
- 适合想提升模型跨文化能力的研究者与开发者。
合成数据已成为扩展大语言模型的核心手段,但其多语言应用仍受限于基于翻译的提示词策略。该方法继承英语中心的表述风格,忽视文化维度,最终制约模型泛化能力。本文认为,被忽略的提示词空间——定义训练分布的关键输入——是提升多语言性能更有力的杠杆。我们提出一种轻量级提示词空间优化框架,对12种语言(涵盖7个语系)的翻译提示词进行自然度、文化适配性与难度增强的系统性转换。在相同数据条件下,相比仅翻译的基线,我们的方法在全局MMLU上提升4.7%,在Flores XCometXL上提升2.4%,在mArenaHard偏好测试中胜出35.3%。结果确立提示词空间优化为构建更鲁棒、文化契合、全球适用的大语言模型的简洁而高效范式。
原文摘要 · Abstract (English)
Synthetic data has become a cornerstone for scaling large language models, yet its multilingual use remains bottlenecked by translation-based prompts. This strategy inherits English-centric framing and style and neglects cultural dimensions, ultimately constraining model generalization. We argue that the overlooked prompt space-the very inputs that define training distributions-offers a more powerful lever for improving multilingual performance. We introduce a lightweight framework for prompt-space optimization, where translated prompts are systematically transformed for Naturalness, Cultural Adaptation, and Difficulty Enhancement. Using an off-the-shelf multilingual LLM, we apply these transformations to prompts for 12 languages spanning 7 families. Under identical data conditions, our approaches achieve substantial and consistent downstream improvements over the translation-only baseline: +4.7% on Global-MMLU accuracy, +2.4% on Flores XCometXL and +35.3% wins in preferences on mArenaHard. We establish prompt-space optimization as a simple yet powerful paradigm for building multilingual LLMs that are more robust, culturally grounded, and globally capable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。