用最优传输理论分析大模型测试时验证的三大核心因素交互
Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimality
- 将验证过程建模为最优传输问题,揭示覆盖率、收敛区与采样偏差的几何关系
- 发现采样偏差-覆盖率曲线存在三种状态:传输、策略改进和饱和
- 适用于研究大模型推理优化、验证机制设计及采样算法选择的研究者
尽管测试时通过验证提升大型语言模型性能已显成效,但验证器的作用及其缺陷仍缺乏深入探索。验证效果由三个量共同决定:(i) 生成器的覆盖率,(ii) 验证器的收敛区域(ROC),(iii) 采样算法的次优性。虽然近期研究部分捕捉了这些因素,但缺乏统一框架来量化它们之间的几何相互作用。本文将可验证的测试时扩展视为一个传输问题,刻画了覆盖率、ROC 与次优性之间的交互,并揭示次优性-覆盖率曲线呈现三种状态:传输阶段(次优性随覆盖率上升)、策略改进阶段(次优性可能随覆盖率下降,取决于验证器的ROC)、饱和阶段(次优性趋于稳定,不受覆盖率影响)。我们进一步提出并分析两类采样算法——顺序与批量,考察其计算复杂度如何影响这些权衡。在 Qwen、Llama 与 Gemma 模型上的实验结果支持了理论发现。
原文摘要 · Abstract (English)
While test-time scaling with verification has shown promise in improving the performance of large language models (LLMs), the role of the verifier and its imperfections remain underexplored. The effect of verification manifests through interactions of three quantities: (i) the generator's coverage, (ii) the verifier's region of convergence (ROC), and (iii) the sampling algorithm's sub-optimality. Though recent studies capture subsets of these factors, a unified framework quantifying the geometry of their interplay is missing. We frame verifiable test-time scaling as a transport problem. This characterizes the interaction of coverage, ROC, and sub-optimality, and uncovers that the sub-optimality--coverage curve exhibits three regimes. A transport regime -- where sub-optimality increases with coverage, a policy improvement regime -- where sub-optimality may decrease with coverage, depending on the verifier's ROC, and a saturation regime -- where sub-optimality plateaus, unaffected by coverage. We further propose and analyze two classes of sampling algorithms -- sequential and batched, and examine how their computational complexities shape these trade-offs. Empirical results with Qwen, Llama, and Gemma models corroborate our theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。