arXiv:2502.19830cs.CLcs.AI2025-02ACL被引 12

通过动态调整温度优化答案聚合,提升大模型推理稳定性。

Revisiting Self-Consistency from Dynamic Distributional Alignment Perspective on Answer Aggregation

  • 基于分布对齐思想,动态调节解码温度以优化采样分布。
  • 在有限样本下优于固定温度方法,平均与最佳性能均提升。
  • 适合需要高效推理的数学题等任务,尤其适用于资源受限场景。

自一致性通过聚合多样化的随机采样结果提升推理能力,但其有效性背后的动态机制仍不明确。本文将自一致性重新建模为一个动态分布对齐问题,揭示解码温度不仅影响采样随机性,更主动塑造潜在答案分布。由于高温度需极大样本量才能稳定,而低温度可能放大偏差,我们提出一种置信度驱动的温度动态校准机制:在不确定性高时锐化采样分布以对齐高概率模式,在置信度高时促进探索。在数学推理任务上的实验表明,该方法在样本有限条件下优于固定多样性基线,提升平均与最优表现,且无需额外数据或模块。这确立了自一致性是采样动态与演化答案分布之间的同步挑战。

原文摘要 · Abstract (English)

Self-consistency improves reasoning by aggregating diverse stochastic samples, yet the dynamics behind its efficacy remain underexplored. We reframe self-consistency as a dynamic distributional alignment problem, revealing that decoding temperature not only governs sampling randomness but also actively shapes the latent answer distribution. Given that high temperatures require prohibitively large sample sizes to stabilize, while low temperatures risk amplifying biases, we propose a confidence-driven mechanism that dynamically calibrates temperature: sharpening the sampling distribution under uncertainty to align with high-probability modes, and promoting exploration when confidence is high. Experiments on mathematical reasoning tasks show this approach outperforms fixed-diversity baselines under limited samples, improving both average and best-case performance across varying initial temperatures without additional data or modules. This establishes self-consistency as a synchronization challenge between sampling dynamics and evolving answer distributions.

大模型推理自一致性动态调参

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。