arXiv:2505.10543cs.AIcs.CL2025-05被引 3

测试大模型在动态任务中的推理能力,发现提示策略能缩小大小模型差距。

Reasoning Capabilities of Large Language Models on Dynamic Tasks

  • 对比自省、启发式变异和规划三种提示策略在动态任务中的表现。
  • 小模型在长提示下表现下降,大模型更稳定;高级提示对小模型提升显著。
  • 尽管提示改进推理,但模型仍缺乏真正涌现的推理能力,尤其在规划与空间协调上。

大语言模型在静态基准上表现出色,但在动态环境中的自学习能力尚不明确。我们评估了三种提示策略——自省、启发式变异和规划——在开放源代码模型上的表现。结果表明,大模型普遍优于小模型,但策略性提示可缩小性能差距。过长的提示会负面影响小模型在基础反应任务中的表现,而大模型则表现更稳健。高级提示技术主要提升小模型在复杂游戏中的表现,对高性能大模型改善有限。然而,先进推理方法效果波动大:当推理与决策一致时可显著提升性能,但也可能引发不稳定甚至大幅性能下降。与人类表现相比,未见真正涌现的推理能力。模型在规划与空间协调方面存在持续局限,表明仅靠自省提示无法完全克服其根本缺陷。推理是多维度任务,尽管链式思维在数学题中有效,但动态基准揭示了通用推理能力的重要短板,提示需超越静态基准以捕捉真实复杂性。

原文摘要 · Abstract (English)

Large language models excel on static benchmarks, but their ability as self-learning agents in dynamic environments remains unclear. We evaluate three prompting strategies: self-reflection, heuristic mutation, and planning across dynamic tasks with open-source models. We find that larger models generally outperform smaller ones, but that strategic prompting can close this performance gap. Second, an overly long prompt can negatively impact smaller models on basic reactive tasks, while larger models show more robust behaviour. Third, advanced prompting techniques primarily benefit smaller models on complex games, but offer less improvement for already high-performing large language models. Yet, we find that advanced reasoning methods yield highly variable outcomes: while capable of significantly improving performance when reasoning and decision-making align, they also introduce instability and can lead to big performance drops. Compared to human performance, our findings reveal little evidence of true emergent reasoning. Instead, large language model performance exhibits persistent limitations in areas like planning and spatial coordination, suggesting that large language models still suffer fundamental shortcomings that may not be fully overcome through self-reflective prompting alone. Reasoning is a multi-faceted task, and while methods like Chain-of-thought improve multi-step reasoning on math word problems, our findings using dynamic benchmarks highlight important shortcomings in general reasoning capabilities, indicating a need to move beyond static benchmarks to capture the complexity of reasoning.

大模型推理能力动态任务提示策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。