arXiv:2505.20184cs.CLcs.AI2025-05被引 3

用思维外显框架评估大模型的高阶思考能力

THiNK: Can Large Language Models Think-aloud?

  • 基于布卢姆分类法设计多智能体反馈评估框架
  • 模型在高阶思维上表现弱,反馈可显著提升表现
  • 适合研究大模型推理与教育认知评估的人看

评估大语言模型的高阶思维能力仍是根本挑战,尤其在超越表面准确性的任务中。本文提出THiNK(Testing Higher-order Notion of Knowledge),一种基于布卢姆分类法的多智能体、反馈驱动的评估框架。该框架将推理评估转化为问题生成、批判与修订的迭代过程,促使大模型通过逐步反思和优化实现思维外显。这使得对低阶(如记忆、理解)和高阶(如评价、创造)思维技能的系统评估成为可能。我们对七种先进大模型应用THiNK,并对其输出进行详细认知分析。结果表明,尽管模型在低阶类别上表现稳定,但在真实情境中应用知识的能力有限,抽象能力较弱。结构化反馈环显著提升了高阶思维表现。定性评估进一步验证,THiNK引导的输出更符合领域逻辑与问题结构。代码已开源,为基于学习科学的评估提供可扩展方法。

原文摘要 · Abstract (English)

Assessing higher-order thinking skills in large language models (LLMs) remains a fundamental challenge, especially in tasks that go beyond surface-level accuracy. In this work, we propose THiNK (Testing Higher-order Notion of Knowledge), a multi-agent, feedback-driven evaluation framework grounded in Bloom's Taxonomy. THiNK frames reasoning assessment as an iterative task of problem generation, critique, and revision, encouraging LLMs to think-aloud through step-by-step reflection and refinement. This enables a systematic evaluation of both lower-order (e.g., remember, understand) and higher-order (e.g., evaluate, create) thinking skills. We apply THiNK to seven state-of-the-art LLMs and perform a detailed cognitive analysis of their outputs. Results reveal that while models reliably perform lower-order categories well, they struggle with applying knowledge in realistic contexts and exhibit limited abstraction. Structured feedback loops significantly improve reasoning performance, particularly in higher-order thinking. Qualitative evaluations further confirm that THiNK-guided outputs better align with domain logic and problem structure. The code of our framework provides a scalable methodology for probing and enhancing LLM reasoning, offering new directions for evaluation grounded in learning science, which is available at our GitHub repository.

大模型评估高阶思维反馈机制认知科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。