提出新评估方法,检测大模型推理稳定性
Are Your LLMs Capable of Stable Reasoning?
- 设计G-Pass@$k$指标,多轮采样评估模型表现与稳定性
- 实验证明现有模型在复杂推理中准确率与一致性不匹配
- 适合关注模型真实推理能力的研究者与应用开发者
大型语言模型(LLMs)在复杂推理任务中取得了显著进展,但基准测试表现与实际应用之间存在明显差距。我们认为这一差距主要源于当前评估协议和指标未能充分捕捉模型在复杂推理任务中的完整能力,尤其是准确性和一致性并重的场景。本文提出G-Pass@$k$,一种新型评估指标,通过多轮采样持续评估模型性能,量化其表现潜力与稳定性。我们在多个公开及新构建的基准上,结合最先进的大语言模型进行大量实验,全面揭示了模型的潜在能力与运行一致性。研究发现,提升大模型真实推理能力仍有巨大空间,凸显了更稳健评估指标的必要性。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) has shown remarkable progress in complex reasoning tasks. However, a significant disparity exists between benchmark performances and real-world applications. We attribute this gap primarily to current evaluation protocols and metrics, which inadequately capture the full spectrum of LLM capabilities, especially in complex reasoning tasks where both accuracy and consistency are essential. In this paper, we introduce G-Pass@$k$, a novel evaluation metric that continuously assesses model performance across multiple sampling attempts, quantifying both the model's performance potential and its stability. Through extensive experiments on various public and newly constructed benchmarks, we employ G-Pass@$k$ in conjunction with state-of-the-art large language models to provide comprehensive insights into their potential capabilities and operational consistency. Our findings reveal a significant opportunity to enhance the realistic reasoning abilities of LLMs, underscoring the necessity for more robust evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。