arXiv:2603.28038cs.AIcs.LG2026-03中稿 · ICLR被引 2

用遗传算法优化提示词,揭示大模型科学推理的隐藏逻辑

Beyond the Answer: Decoding the Behavior of LLMs as Scientific Reasoners

  • 用改进的遗传帕累托算法系统优化科学推理提示词
  • 发现模型推理依赖特定逻辑,跨模型难以通用
  • 适合研究大模型可解释性与人机协作的学者

随着大语言模型在复杂推理任务中表现日益出色,现有架构成为前沿模型内部启发式策略的重要代理。揭示涌现推理机制对长期可解释性和安全性至关重要。此外,理解提示如何调控这些过程也极为关键,因为自然语言很可能是与未来通用人工智能系统交互的主要方式。本文采用自定义的遗传帕累托(GEPA)算法,系统优化科学推理任务的提示词,并分析提示对推理行为的影响。我们研究了GEPA优化提示中的结构模式与逻辑启发式,并评估其可迁移性与脆弱性。结果表明,科学推理性能的提升往往对应于模型特有的启发式策略,这些策略在不同系统间无法泛化,我们称之为‘局部’逻辑。通过将提示优化视为模型可解释性的工具,我们认为绘制大模型偏好的推理结构是有效协同超智能系统的重要前提。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) achieve increasingly sophisticated performance on complex reasoning tasks, current architectures serve as critical proxies for the internal heuristics of frontier models. Characterizing emergent reasoning is vital for long-term interpretability and safety. Furthermore, understanding how prompting modulates these processes is essential, as natural language will likely be the primary interface for interacting with AGI systems. In this work, we use a custom variant of Genetic Pareto (GEPA) to systematically optimize prompts for scientific reasoning tasks, and analyze how prompting can affect reasoning behavior. We investigate the structural patterns and logical heuristics inherent in GEPA-optimized prompts, and evaluate their transferability and brittleness. Our findings reveal that gains in scientific reasoning often correspond to model-specific heuristics that fail to generalize across systems, which we call "local" logic. By framing prompt optimization as a tool for model interpretability, we argue that mapping these preferred reasoning structures for LLMs is an important prerequisite for effectively collaborating with superhuman intelligence.

大模型推理提示优化可解释性科学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。