用模型失败反向生成动态测试题,更真实评估大模型推理能力
StressEval: Failure-Driven Dynamic Benchmarking for Knowledge-Intensive Reasoning in Large Language Models

- 基于模型错误自动构造挑战性测试题
- 在多个SOTA模型上引发显著性能下降
- 适合评估和改进知识推理型大模型
静态基准因数据污染和过拟合正日益失效,尤其在知识密集型推理任务中。现有动态基准虽缓解了陈旧问题,却常以牺牲答案可得性和可控性为代价。本文提出StressEval——一种故障驱动的数据合成框架,将模型失败转化为动态、挑战性强且可控的测试实例。该框架包含三个阶段:首先构建半结构化的难度卡片,识别失败的推理步骤及其根本原因;其次采用双视角实例生成方法,同时针对知识缺口与推理断裂点进行优化,同时保留底层难度因素;最后通过门控机制仅保留有依据、无歧义的实例。基于多个知识密集型推理数据集,我们用StressEval构建了动态基准套件Dynamic OneEval。在多个先进大模型上,Dynamic OneEval展现出比原始基准更显著的性能下降,同时保留明确的难度因素,支持更具针对性的迭代改进。
原文摘要 · Abstract (English)
Static benchmarks for LLMs are increasingly compromised by contamination and overfitting especially on knowledge intensive reasoning tasks While recent dynamic benchmarks can alleviate staleness they often increase difficulty at the expense of answerability and controllability In this paper we propose StressEval a failure driven data synthesis framework that turns observed model failures into dynamic challenging and controllable test instances StressEval consists of three stages first it constructs a semi structured difficulty card that identifies the failed reasoning step and its root cause second it applies a dual perspective instance synthesis method that targets both knowledge gaps and reasoning breakdowns while preserving the underlying difficulty factors and third it applies a gating mechanism to retain only grounded unambiguous instances Seeding from multiple knowledge intensive reasoning datasets we employ StressEval to build Dynamic OneEval a focused suite of challenging dynamic benchmark Across several state of the art LLMs Dynamic OneEval yields substantially larger performance drops than the original benchmarks while retaining explicit difficulty factors enabling more actionable iteration
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。