arXiv:2508.16741cs.LGcs.AI2025-08

小模型通过强化学习帮大模型优化提示词,提升性能且更安全。

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning

  • 用小模型生成指令,通过强化学习不断优化,提升大模型表现。
  • 在MATH-500上达到98%准确率,HH-RLHF上提升134%。
  • 适合无法微调大模型或需安全提示的场景,高效可靠。

有效的提示工程对许多应用仍具挑战性。我们提出弱到强知识迁移(WST),一种自动提示工程框架,其中小型“教师”模型生成能提升大型“学生”模型性能的指令。与以往工作不同,WST仅需弱教师模型,使其在大模型闭源或难以微调的场景中依然高效且普适。通过强化学习,教师模型的指令根据学生模型的输出迭代优化,在推理(MATH-500、GSM8K)和对齐(HH-RLHF)基准上取得显著提升——MATH-500达98%,HH-RLHF提升134%,超越GPT-4o-mini和Llama-70B等基线。结果表明,小模型可稳定引导大模型释放潜在能力,避免强教师引入误导性提示,确立WST为高效、安全的大语言模型提示优化可扩展方案。

原文摘要 · Abstract (English)

Effective prompt engineering remains a challenging task for many applications. We introduce Weak-to-Strong Transfer (WST), an automatic prompt engineering framework where a small "Teacher" model generates instructions that enhance the performance of a much larger "Student" model. Unlike prior work, WST requires only a weak teacher, making it efficient and broadly applicable in settings where large models are closed-source or difficult to fine-tune. Using reinforcement learning, the Teacher Model's instructions are iteratively improved based on the Student Model's outcomes, yielding substantial gains across reasoning (MATH-500, GSM8K) and alignment (HH-RLHF) benchmarks - 98% on MATH-500 and 134% on HH-RLHF - and surpassing baselines such as GPT-4o-mini and Llama-70B. These results demonstrate that small models can reliably scaffold larger ones, unlocking latent capabilities while avoiding misleading prompts that stronger teachers may introduce, establishing WST as a scalable solution for efficient and safe LLM prompt refinement.

提示工程强化学习小模型引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。