通过多样化采样提升大模型推理性能,理论揭示其增效机制。
On the Effect of Sampling Diversity in Scaling LLM Inference
- 用多样化提示生成响应,显著降低错误率。
- 发现多样性和精度存在权衡,适配不同任务可获更强效果。
- 适合关注推理优化与模型效率的研究者参考。
大语言模型(LLM)的推理缩放是提升性能的关键,而引入多样性已被证明是有效手段。基于解的准确率与有意义响应多样性之间的观察关系,我们系统研究了提示多样性在推理缩放中的作用。理论上解释为何多样化采样能改善Best-of-$N$缩放,表明经过Best-of-$N$选择后,从多样化提示生成的响应比静态提示生成的响应误差率显著更低。基于此分析,我们推导出多样性-保真度权衡原则,指导多样化采样策略的设计。由此衍生出一系列有效的扰动风格。我们理论和实证地刻画了多样化探索何时仍有效,表明其在多种条件下均成立,并进一步指出在多数投票机制下,多样性可能消失。最后,我们系统评估了采样多样性的有效性,显示在适当场景下,随着多样性增加,有意义的扰动能带来更强、任务依赖的性能提升。整体上,该工作提供了理解采样多样性如何影响LLM推理时缩放的理论与实证基础。
原文摘要 · Abstract (English)
Large language model (LLM) scaling inference is key to unlocking greater performance, and leveraging diversity has proven an effective way to enhance it. Motivated by the observed relationship between solution accuracy and meaningful response diversity, we systematically study the effect of prompt diversity in scaling inference. We theoretically explain why diversified sampling improves Best-of-$N$ scaling, showing that responses generated from diverse prompts after Best-of-$N$ selection exhibit significantly lower error rates than those produced from stationary prompts. Building on this analysis, we derive a diversity-fidelity trade-off principle, that guides the design of sampling strategies introducing diversity. From this guidance, we instantiate a family of effective perturbation styles. We theoretically and empirically characterize \textbf{when} diversified exploration remains effective, demonstrating that it works under a variety of conditions, and we further show that under majority voting, diversity may vanish. Finally, we systematically evaluate the effectiveness of sampling diversity and show that, when applied appropriately in different contexts, meaningful perturbations yield stronger, task-dependent gains as diversity increases. Overall, this work provides a systematic analysis that offers a theoretical and empirical foundation for understanding how sampling diversity affects LLM inference-time scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。