为大模型推理提供可证明的可靠性保障,统一解释自一致与测试时强化机制。
Certified Self-Consistency: Statistical Guarantees and Test-Time Training for Reliable Reasoning in LLMs
- 基于多数投票构建统计证书,证明多数答案等于模型输出的众数。
- 提出马尔可夫多数证书(MMC),可动态判断采样是否足够,保证置信度。
- 揭示无标签训练能通过指数倾斜分布减少采样量,适合追求可靠推理的研究者。
近期的自一致和测试时强化学习(TTRL)等方法在无需额外监督的情况下提升了大语言模型(LLMs)的推理可靠性,但其内在机制和统计保证仍不清晰。本文提出一个统一的可证推理框架,表明多数投票在弱假设下能以高概率捕获模型最终输出分布的众数。我们推导了有限样本和任意时间有效的集中界,量化了这一置信度,并引入马尔可夫多数证书(MMC),一种自适应的序列停止规则,用于确定何时已采集足够样本。进一步证明,诸如TTRL等无标签后训练方法通过指数倾斜答案分布向其众数靠拢,从而降低认证所需样本数。基于此洞察,我们提出了新的后训练目标,显式优化尖锐性与偏差之间的权衡。这些结果在一个统一的统计框架下解释并连接了两种核心的测试时缩放策略:自一致与TTRL,实现了无标签、可证可靠的推理。
原文摘要 · Abstract (English)
Recent advances such as self-consistency and test-time reinforcement learning (TTRL) improve the reliability of large language models (LLMs) without additional supervision, yet their underlying mechanisms and statistical guarantees remain poorly understood. We present a unified framework for certifiable inference in LLMs, showing that majority voting provides a statistical certificate of self-consistency: under mild assumptions, the aggregated answer coincides with the mode of the model's terminal distribution with high probability. We derive finite-sample and anytime-valid concentration bounds that quantify this confidence, and introduce the Martingale Majority Certificate (MMC), a sequential stopping rule that adaptively determines when sufficient samples have been drawn. We further prove that label-free post-training methods such as TTRL implicitly sharpen the answer distribution by exponentially tilting it toward its mode, thereby reducing the number of samples required for certification. Building on this insight, we propose new post-training objectives that explicitly optimise this trade-off between sharpness and bias. Together, these results explain and connect two central test-time scaling strategies, self-consistency and TTRL, within a single statistical framework for label-free, certifiable reliability in reasoning LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。