优化系统提示词,让大模型跨语言推理更准确可靠。
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
- 构建四维评估框架,量化分析提示词在多语言中的表现。
- 自动优化提示词可提升各指标5%-10%,减少语言混杂。
- 优质提示词促进更结构化推理,降低无效语言切换。
系统提示词是一种轻量但强大的推理时控制大语言模型(LLMs)的方法。尽管以往研究集中于英语环境,真实场景中需统一提示实现跨语言稳定运行。本文对不同系统提示如何引导模型实现准确、鲁棒的跨语言行为进行了全面研究。提出一个统一的四维评估框架,在五种语言、三类大模型和三个基准上开展大规模实验,发现思维链(CoT)、情绪、场景等提示成分与多语言鲁棒性正相关。开发了面向多语言的提示优化框架,可自动发现使各项指标提升5%-10%的提示。分析超1000万推理单元后发现,性能更强的提示能诱导更结构化、一致的推理模式,并减少不必要的语言切换。结果表明,系统提示优化是实现高效跨语言大模型行为的可行路径。
原文摘要 · Abstract (English)
System prompts provide a lightweight yet powerful mechanism for conditioning large language models (LLMs) at inference time. While prior work has focused on English-only settings, real-world deployments benefit from having a single prompt to operate reliably across languages. This paper presents a comprehensive study of how different system prompts steer models toward accurate and robust cross-lingual behavior. We propose a unified four-dimensional evaluation framework to assess system prompts in multilingual environments. Through large-scale experiments on five languages, three LLMs, and three benchmarks, we uncover that certain prompt components, such as CoT, emotion, and scenario, correlate with robust multilingual behavior. We develop a prompt optimization framework for multilingual settings and show it can automatically discover prompts that improve all metrics by 5-10%. Finally, we analyze over 10 million reasoning units and find that more performant system prompts induce more structured and consistent reasoning patterns, while reducing unnecessary language-switching. Together, we highlight system prompt optimization as a scalable path to accurate and robust multilingual LLM behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。