arXiv:2507.23776cs.CL2025-07

通过逐步披露问题信息,更真实评估大模型的解题能力。

Cascaded Information Disclosure for Generalized Evaluation of Problem Solving Capabilities

  • 分阶段逐步揭示问题线索,激发模型通用推理能力。
  • 在多个数据集上缩小了不同模型间的性能差距。
  • 适合想深入评估模型真实推理水平的研究者。

尽管问答(QA)基准测试是评估大语言模型性能的一种自动且可扩展的方法,但其仅间接反映模型的底层解题能力。为此,我们提出一种基于级联式问题披露(cascaded question disclosure)的综合、可泛化框架,可在保持自动化与可扩展性的同时,更准确地估计模型的解题能力。该方法以分阶段方式收集模型响应,每阶段逐步披露问题的部分信息,旨在激发模型的通用推理行为。实验表明,该方法不仅提升了模型间比较的准确性,还促使模型产生更优的中间推理轨迹,相较于标准QA范式表现更佳。我们在多种推理与知识密集型的QA数据集上,对比了不同规模和架构的大模型,验证了该方法的有效性。结果发现,该方法缩小了标准QA评价中观察到的性能差距,说明现有间接评估方式高估了模型间的实际差异。进一步的消融实验也验证了该结论的稳健性。

原文摘要 · Abstract (English)

While question-answering~(QA) benchmark performance is an automatic and scalable method to compare LLMs, it is an indirect method of evaluating their underlying problem-solving capabilities. Therefore, we propose a holistic and generalizable framework based on \emph{cascaded question disclosure} that provides a more accurate estimate of the models' problem-solving capabilities while maintaining the scalability and automation. This approach collects model responses in a stagewise manner with each stage revealing partial information about the question designed to elicit generalized reasoning in LLMs. We find that our approach not only provides a better comparison between LLMs, but also induces better intermediate traces in models compared to the standard QA paradigm. We empirically verify this behavior on diverse reasoning and knowledge-heavy QA datasets by comparing LLMs of varying sizes and families. Our approach narrows the performance gap observed in the standard QA evaluation settings, indicating that the prevalent indirect QA paradigm of evaluation overestimates the differences in performance between models. We further validate our findings by extensive ablation studies.

大模型评估推理能力级联披露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。