让大模型通过反思不一致答案提升推理准确性和可信度
Mirror-Consistency: Harnessing Inconsistency in Majority Voting
- 引入'反射镜'机制,分析多个生成结果间的不一致
- 在数学和常识推理任务上准确率提升2.3%~5.1%
- 适合需要高可靠性的决策类AI应用
自一致性(Self-Consistency)是一种广泛使用的解码策略,显著提升了大语言模型(LLM)的推理能力。然而,它依赖于多数投票规则,仅关注最频繁的答案,忽视了其他少数观点。这些不一致的少数意见往往揭示了模型生成过程中的不确定性。为解决这一局限,我们提出镜像一致性(Mirror-Consistency),作为标准自一致性方法的增强。该方法在自集成解码过程中引入一个‘反射镜’机制,使模型能够批判性地审视多个生成结果之间的不一致。此外,正如人类利用镜子更好地认识自我,我们提出使用镜像一致性来改进基于样本的置信度校准方法,有助于缓解模型过度自信的问题。实验结果表明,与自一致性相比,镜像一致性在推理准确性和置信度校准方面均表现出更优性能。
原文摘要 · Abstract (English)
Self-Consistency, a widely-used decoding strategy, significantly boosts the reasoning capabilities of Large Language Models (LLMs). However, it depends on the plurality voting rule, which focuses on the most frequent answer while overlooking all other minority responses. These inconsistent minority views often illuminate areas of uncertainty within the model's generation process. To address this limitation, we present Mirror-Consistency, an enhancement of the standard Self-Consistency approach. Our method incorporates a 'reflective mirror' into the self-ensemble decoding process and enables LLMs to critically examine inconsistencies among multiple generations. Additionally, just as humans use the mirror to better understand themselves, we propose using Mirror-Consistency to enhance the sample-based confidence calibration methods, which helps to mitigate issues of overconfidence. Our experimental results demonstrate that Mirror-Consistency yields superior performance in both reasoning accuracy and confidence calibration compared to Self-Consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。