arXiv:2605.28465cs.CL2026-05

评测并提升大模型在交互中的发散思维能力,发现其易陷入固定动作陷阱。

Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents

论文配图:Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents
图 1 · 摘自论文原文
  • 设计交互式评测基准MUTATE,分路径与动作两个层级评估发散思维
  • 前沿模型在压力下易固定动作,仅18%路径实现动作级发散
  • 提出ReDNA方法,分离发散生成与收敛选择,显著提升多路径探索能力

发散思维是创造力的核心维度,但现有大语言模型(LLMs)评估多为单轮文本生成,无法捕捉代理在迭代交互中的推理过程。为此,我们提出MUTATE,一个交互式基准,用于评估代理在路径级(发现达成同一目标的多种路径)和动作级(使用非典型方式操作对象)两个层面的发散思维能力。不同于仅计成功路径的评估,MUTATE同时评分完成路径与未完成尝试,保留传统成功率忽略的发散推理。对前沿LLMs的实验显示,现有框架存在结构性盲点:在即时收敛压力下,模型容易陷入立即行动固化,动作级发散率不足18%。为解决此问题,我们提出ReDNA,将不受限的发散候选生成与受约束的收敛选择相分离。ReDNA在两个发散层级均显著优于先前方法,并在外源创造力环境中有良好泛化能力。我们进一步验证其成功源于韧性发散推理的质变,而非简单的环境探索。

原文摘要 · Abstract (English)

Divergent thinking is a core dimension of creativity, yet existing evaluations of Large Language Models (LLMs) treat them as single-turn text generations, failing to capture how an agent reasons through iterative interaction. To address this, we introduce MUTATE, an interactive benchmark designed to evaluate agentic divergent thinking at two levels: path-level, where an agent discovers multiple alternative paths to the same goal, and action-level, where individual actions require non-typical, mechanism-shifting object uses. Unlike success-only evaluations, MUTATE scores both completed paths and off-path attempts, capturing divergent reasoning that conventional success rates discard. Our experiments with frontier LLMs reveal a structural blind spot in existing frameworks: when exposed to immediate convergence pressure, they tend to fall into immediate action fixation, failing to improve action-level divergence. To overcome this, we propose ReDNA, which separates unconstrained divergent candidate generation from convergent constraint selection. ReDNA significantly outperforms prior methods across both divergence levels and generalizes effectively to an external creativity environment. We also confirm its success stems from a qualitative enhancement of resilient divergent reasoning rather than simple environmental exploration.

发散思维交互智能体大模型评测创造力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。