arXiv:2510.10528cs.CLcs.LG2025-10ACL

用黑箱说服提示让大模型少思考却不失准,显著降低推理成本。

Merlin's Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting

  • 通过多角度迭代生成说服性提示,引导模型精简输出。
  • 在GSM8K上使通义千问响应长度减少3倍,整体平均降40%耗能。
  • 适配不同模型和数据域,对闭源API也有效,实用性强。

大型推理模型(LRMs)虽能通过逐步思考完成复杂任务,但其长推理过程带来巨大计算与延迟开销,阻碍实际部署。本文提出一种基于黑箱说服性提示的新方法,将LRM视为黑箱通信器,研究如何说服其生成简洁回答而不牺牲准确率。我们设计了Whisper框架,从多元视角迭代生成高质量说服性提示。跨多个基准的实验表明,Whisper持续减少令牌使用量,同时保持性能。尤其在简单GSM8K问题上,通义千问系列模型平均响应长度减少3倍;所有基准平均节省约40%令牌。对于闭源API,在MATH-500上,Claude-3.7和Gemini-2.5分别减少46%和50%令牌消耗。进一步分析显示,Whisper在不同数据领域、模型规模和模型家族中均具广泛适用性,验证了黑箱说服性提示作为提升LRM效率的实用策略潜力。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) have demonstrated remarkable proficiency in tackling complex tasks through step-by-step thinking. However, this lengthy reasoning process incurs substantial computational and latency overheads, hindering the practical deployment of LRMs. This work presents a new approach to mitigating overthinking in LRMs via black-box persuasive prompting. By treating LRMs as black-box communicators, we investigate how to persuade them to generate concise responses without compromising accuracy. We introduce Whisper, an iterative refinement framework that generates high-quality persuasive prompts from diverse perspectives. Experiments across multiple benchmarks demonstrate that Whisper consistently reduces token usage while preserving performance. Notably, Whisper achieves a 3x reduction in average response length on simple GSM8K questions for the Qwen3 model series and delivers an average ~40% token reduction across all benchmarks. For closed-source APIs, Whisper reduces token usage on MATH-500 by 46% for Claude-3.7 and 50% for Gemini-2.5. Further analysis reveals the broad applicability of Whisper across data domains, model scales, and families, underscoring the potential of black-box persuasive prompting as a practical strategy for enhancing LRM efficiency.

推理优化提示工程大模型效率黑箱策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。