LLM推理结果不稳定,单次评分会误判系统性能。
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
- 用30次独立实验捕捉推理质量与成本的分布特性
- 77%胜率表明同一策略在对抗中常被压过一头
- 提出新评估范式:关注分布而非单一得分
大型语言模型(LLM)推理系统的评测分数常以单一数值呈现,但相同模型、策略和任务在重复执行中仍会产生显著不同的答案与成本,即使使用贪婪解码(T=0)亦然。这种波动并非随机噪声:表现最佳的策略在与最接近的对手对决时仅赢77%的比赛,说明单次得分可能错误排名系统。我们提出ReasonBench基准套件,记录12个模型、10种推理策略、6个任务下的30次独立实验,将质量与成本视为分布而非点估计。研究发现,该波动具有结构性:通过全局噪声(跨基准不均)与运行噪声(组内随机性)的两分法可揭示策略架构对稳定性的影响;模型与策略分别影响分布的不同维度。层级分解显示,75%的分数方差源于基准、系统与题目结构,另有持续残留项被单次评估所掩盖。此外,成本与质量非对称解耦:低成本方法在成本-质量联合失败上具有结构性免疫,而高成本方法无论准确性如何始终暴露。这些发现确立了不稳定性是推理系统固有属性,推动分布感知评估成为标准实践。
原文摘要 · Abstract (English)
Benchmark scores for LLM reasoning systems are reported as single numbers, yet the same model, strategy, and task can produce meaningfully different answers and costs across repeated executions, even under greedy decoding (T = 0). This variance is not a statistical nuisance: the highest-performing strategy wins only 77% of head-to-head runs against its nearest competitor, meaning a single observed score can silently misrank systems. We introduce ReasonBench, a benchmark suite recording 30 independent trials across 10 reasoning strategies, 12 models, and 6 tasks, treating quality and cost as distributions rather than point estimates. We find that this variance is structured rather than random: a two-component taxonomy -- Global Noise, capturing cross-benchmark unevenness, and Run Noise, capturing within-benchmark stochasticity -- reveals that strategy architecture predicts stability profiles, while models and strategies shift orthogonal aspects of the distribution. A hierarchical decomposition attributes three-quarters of score variance to benchmark, system, and item structure, with a persistent residual that single-run evaluation silently absorbs. Finally, cost and quality decouple asymmetrically: cheap methods are structurally immune to joint cost-quality failure, while expensive methods remain exposed regardless of their accuracy. These findings establish instability as an inherent property of reasoning systems and motivate distribution-aware evaluation as standard practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。