arXiv:2604.03374cs.CLcs.AI2026-04中稿 · COLM

构建真实世界知识的创意解题基准,检验模型跨领域联想能力

CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge

  • 基于真实世界知识设计五级难度谜题
  • 模型在创意题上表现比事实题差17%
  • 适合评估大模型跨域联想与创造性推理能力

创意问题解决需要逻辑推理、横向思维、类比和常识知识等多种认知能力的结合,以发现看似无关信息间的深层关联。然而,现有大语言模型评测多仅关注单一环节,且多数创意类基准依赖人为构造的脑筋急转弯或虚构情境,无法反映真实世界中的创意过程。为此,我们提出CresOWLve(含2,061个实例,分五个难度等级),一个基于真实世界知识的创意解题评测基准。其题目要求综合运用多种创造性思维策略,从多元领域检索事实并创造性整合以得出答案。对前沿非思考型与思考型模型的评估显示,CresOWLve仍具高度挑战性。分析表明:模型在事实类问题上表现显著优于创意类问题(最高下降17%)。尽管能准确检索相关知识,但难以形成非显性的创造性连接以得出正确答案。

原文摘要 · Abstract (English)

Creative problem-solving requires combining multiple cognitive abilities, including logical reasoning, lateral thinking, analogy-making, and commonsense knowledge, to discover insights that connect seemingly unrelated pieces of information. However, most existing benchmarks for large language models (LLMs) evaluate only specific components of this process. Moreover, many creativity-oriented benchmarks rely on artificially constructed brainteasers or contrived scenarios that do not reflect how creative problem-solving occurs in real-world settings. To address this gap, we introduce CresOWLve (containing 2,061 examples across five difficulty levels), a benchmark for evaluating creative problem-solving using puzzles grounded in real-world knowledge. Problems in CresOWLve require employing multiple creative thinking strategies, retrieving facts from diverse domains, and creatively combining them to arrive at a solution. Evaluating several frontier non-thinking and thinking LLMs, we show that CresOWLve remains highly challenging. Our analysis reveals a consistent performance gap: models perform substantially better on factual questions than on creative ones (up to -17% drop). While models can often retrieve the relevant knowledge, they struggle to form the non-obvious creative connections required to integrate the knowledge and arrive at the correct answer.

创意推理知识融合评测基准LLM评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。