arXiv:2505.24532cs.CL2025-05被引 1

用认知分级框架生成真实世界难题,揭示大模型推理能力被高估

DeepQuestion: Systematic Generation of Real-World Challenges for Evaluating LLMs Performance

  • 基于布卢姆分类法自动化提升数据集认知难度
  • 模型准确率最高下降70%,显示真实场景下能力严重下滑
  • 适合评估大模型真实推理能力的研究者与开发者

尽管大语言模型在标准基准上接近人类表现,其能力却难以泛化到复杂的现实问题。为弥合这一差距,我们提出DeepQuestion——一个可扩展的自动化框架,系统性提升现有数据集的认知复杂度。该框架基于布卢姆分类法,生成两类任务:(1) 情景式问题,测试模型在嘈杂、真实情境中应用知识的能力;(2) 基于指令的提示,要求模型从给定解题路径生成新问题,评估其综合与评价能力。我们在十种主流开源及专有模型上进行广泛评估,结果显示,随着任务认知层级上升,模型准确率最高下降70%。这些发现表明,当前基准严重夸大了模型的真实推理能力,凸显了开展认知多样性评估对指导未来大模型发展的关键意义。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) achieve near-human performance on standard benchmarks, their capabilities often fail to generalize to complex, real-world problems. To bridge this gap, we introduce DeepQuestion, a scalable, automated framework that systematically elevates the cognitive complexity of existing datasets. Grounded in Bloom's taxonomy, DeepQuestion generates (1) scenario-based problems to test the application of knowledge in noisy, realistic contexts, and (2) instruction-based prompts that require models to create new questions from a given solution path, assessing synthesis and evaluation skills. Our extensive evaluation across ten leading open-source and proprietary models reveals a stark performance decline with accuracy dropping by up to 70% as tasks ascend the cognitive hierarchy. These findings underscore that current benchmarks overestimate true reasoning abilities and highlight the critical need for cognitively diverse evaluations to guide future LLM development.

大模型评估认知层级真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。