arXiv:2603.27844cs.CL2026-03被引 2

模型能力比提示工程更重要,多模型投票效果受限于错误相关性。

Model Capability Dominates: Inference-Time Optimization Lessons from AIMO 3

  • 用不同推理策略分配给多个模型,尝试降低错误相关性
  • 8个模型能力差距下,能力强者表现远超提示优化效果
  • 提升准确率需依赖模型自身能力,而非提示设计

对多个大语言模型的推理结果进行多数投票可提升数学推理能力,但错误相关性限制了有效样本量。一个自然的解决方案是为不同模型分配不同的推理策略。该方法在AIMO 3竞赛中验证:使用3个模型,23+次实验,50道国际数学奥林匹克级别问题,单块H100 80GB显卡,5小时时间限制。所有提示层级干预均失败。高温采样已能有效解相关;较弱策略反而降低准确率且未显著减少相关性。在等样本数N=8的情况下,8点能力差距下模型能力始终占优。最佳多数投票得分(42/50)与pass@20(约45.5)之间的差距是选择损失,非提示损失。验证器驱动的选择器可弥补,但提示工程无法做到。

原文摘要 · Abstract (English)

Majority voting over multiple LLM attempts improves mathematical reasoning, but correlated errors limit the effective sample size. A natural fix is to assign different reasoning strategies to different voters. The approach, Diverse Prompt Mixer, is tested on the AIMO 3 competition: 3 models, 23+ experiments, 50 IMO-level problems, one H100 80 GB, 5-hour limit. Every prompt-level intervention fails. High-temperature sampling already decorrelates errors; weaker strategies reduce accuracy more than they reduce correlation. Across an 8-point capability gap at equal N=8 and every optimization tested, model capability dominates. The gap between the best majority-vote score (42/50) and pass@20 (~45.5) is selection loss, not prompt loss. A verifier-based selector could close it. Prompt engineering cannot.

大模型推理多模型投票能力主导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。