arXiv:2510.02377cs.CLcs.LG2025-10EMNLP被引 7

用置信度评分从多个大模型中选最优答案,提升推理准确性。

Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems

  • 基于校准后对数似然得分,自动评估多模型输出可靠性。
  • 在GSM8K、MMLU等数据集上准确率提升4%~5%。
  • 无需外部验证或多次采样,适合资源受限场景。

大型语言模型(LLMs)展现出卓越能力,但在资源受限环境下,从多个模型中选出最可靠的回应仍具挑战。现有方法常依赖昂贵的外部验证器、人工评估或需单模型多次采样的自一致技术。尽管多模型系统比单模型生成更丰富的响应,潜力更大,但性能往往不及单模型自一致方法。本文提出一种原理严谨、新颖且计算高效的多模型响应选择方法,利用校准后的对数似然分数,隐式融合各模型的内在知识与置信度。该方法在GSM8K、MMLU(6个子集)和ARC数据集上,于辩论(多轮模型讨论)和非辩论(多模型Best-of-N)设置下分别实现约4%、3%和5%的性能提升。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated exceptional capabilities, yet selecting the most reliable response from multiple LLMs remains a challenge, particularly in resource-constrained settings. Existing approaches often depend on costly external verifiers, human evaluators, or self-consistency techniques that require multiple samples from a single model. While multi-LLM systems produce more diverse responses than single models and thus have greater potential, they often underperform compared to single LLM self-consistency. We propose a principled, novel and computationally efficient method to select the best response from multiple different LLMs using a calibrated log-likelihood score, implicitly leveraging the inherent knowledge and confidence of these models. Our method demonstrates improvements of approx. 4%, 3%, and 5% across both debate (multi-round LLM discussions) and non-debate (Best-of-N with multiple LLMs) settings on GSM8K, MMLU (6 subsets), and ARC datasets respectively.

多模型答案选择推理增强置信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。