arXiv:2605.02910cs.AIcs.CL2026-05被引 2

评测大模型通过物体属性创新使用工具的能力,发现现有模型创意推理仍严重不足。

CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing

论文配图:CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing
图 1 · 摘自论文原文
  • 基于物体属性构建4000+实体的知识库,生成1.4万道需创造性解题的任务
  • 10个顶尖模型在任务中仅能正确选物,难识别关键部件与物理机制
  • 模型规模提升难以突破瓶颈,链式思考等策略效果有限,适合研究智能体创意

大语言模型在推理和环境交互任务上表现优异,但其创造性问题解决能力仍待探索。本文从创造性工具使用视角出发,考察模型通过推理物体功能与属性进行非标准用途改造的能力。为此,我们构建了包含4000个实体和15万+属性标注的大型功能知识库(KB),明确关联物体、部件、属性与可操作用途。基于该知识库,生成14,000个需在约束下提出非显而易见但物理可行解的实证任务。对10个主流大模型(含闭源与开源)的评估显示:模型虽常能选出合理物体,却普遍无法准确识别关键部件、其功能属性及背后物理机制,导致性能显著下降。此外,模型规模扩大带来的改进迅速饱和,强通用推理能力并不保证创造性功能发现,常用推理策略如链式思考(Chain-of-Thought)增益有限。结果表明,创造性工具使用仍是当前模型的重大挑战,CreativityBench为研究这一智能缺失维度提供了有效测试平台,对未来智能体规划与推理模块具有潜在启示。

原文摘要 · Abstract (English)

Recent advances in large language models have led to strong performance on reasoning and environment-interaction tasks, yet their ability for creative problem-solving remains underexplored. We study this capability through the lens of creative tool use, where a model repurposes available objects by reasoning about their affordances and attributes rather than relying on canonical usage. As a first step, we introduce CreativityBench, a benchmark for evaluating affordance-based creativity in LLMs. To this end, we build a large-scale affordance knowledge base (KB) with 4K entities and 150K+ affordance annotations, explicitly linking objects, parts, attributes, and actionable uses. Building on this KB, we generate 14K grounded tasks that require identifying non-obvious yet physically plausible solutions under constraints. Evaluations across 10 state-of-the-art LLMs, including closed and open-source models, show that models can often select a plausible object, but fail to identify the correct parts, their affordances, and the underlying physical mechanism needed to solve the task, leading to a significant drop in performance. Furthermore, improvements from model scaling quickly saturate, strong general reasoning does not reliably translate to creative affordance discovery, and common inference-time strategies such as Chain-of-Thought yield limited gains. These results suggest that creative tool use remains a major challenge for current models, and that CreativityBench provides a useful testbed for studying this missing dimension of intelligence, with potential implications for planning and reasoning modules in future agents.

创意推理工具使用大模型评测知识库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。