arXiv:2504.00762cs.AI2025-04AAAI被引 23

用多个模型轮流生成并投票,省钱又提效。

Do We Truly Need So Many Samples? Multi-LLM Repeated Sampling Efficiently Scales Test-Time Compute

  • 多模型交替生成,靠一致性自动选最优
  • 六数据集测试,性能超自洽和多智能体辩论
  • 只需少量模型,适合资源有限的部署场景

本文提出一种简单高效且低成本的策略,通过扩展测试时计算能力来提升大语言模型性能。该策略基于重复采样后投票框架,创新性地引入多个不同模型(包括较弱模型),利用它们因训练数据与范式差异带来的互补优势。通过一致性信号动态切换模型,理论分析表明该方法兼具效率与性能优势。在六个数据集上的实验显示,该方法不仅优于自洽性及当前最先进的多智能体辩论方法,还显著降低推理成本。此外,ModelSwitch仅需少数相近水平的LLM即可达到最优表现,并可结合验证机制扩展,展现出在生成-验证范式中利用多模型的巨大潜力。

原文摘要 · Abstract (English)

This paper presents a simple, effective, and cost-efficient strategy to improve LLM performance by scaling test-time compute. Our strategy builds upon the repeated-sampling-then-voting framework, with a novel twist: incorporating multiple models, even weaker ones, to leverage their complementary strengths that potentially arise from diverse training data and paradigms. By using consistency as a signal, our strategy dynamically switches between models. Theoretical analysis highlights the efficiency and performance advantages of our strategy. Extensive experiments on six datasets demonstrate that our strategy not only outperforms self-consistency and state-of-the-art multi-agent debate approaches, but also significantly reduces inference costs. Additionally, ModelSwitch requires only a few comparable LLMs to achieve optimal performance and can be extended with verification methods, demonstrating the potential of leveraging multiple LLMs in the generation-verification paradigm.

大模型测试时计算多模型协作推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。