arXiv:2605.26655cs.CLcs.LG2026-05

分析提示优化为何有效或失效,发现不同编辑类型对任务表现有系统性影响。

Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis

论文配图:Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis
图 1 · 摘自论文原文
  • 通过因果启发的编辑级分析,识别出提示优化中的模式规律。
  • 增加复杂度和元指令会降低数学与多跳推理性能,而分步与元认知编辑提升逻辑推理。
  • 结果跨框架、模型和任务通用,为定制化优化器设计提供依据。

自动化提示优化方法(如DSpy、TextGrad)可显著提升大语言模型(LLM)性能,但其在不同任务间的泛化能力仍不理想。实践中,优化提示在某一基准上的优势常无法迁移到另一任务,且这一局限性在更换不同LLM骨干模型后依然存在。为探究提示性能异质性的潜在根源,本文基于因果推断思想,对多种优化框架、LLM骨干模型及自然语言处理基准上的优化提示进行观察性分析。通过倾向性调整的关联分析结合多种提示编辑表征,识别出一致的任务条件编辑模式。研究发现:复杂度增加与元指令类编辑与数学及多跳推理性能呈负相关;而分步式与元认知编辑则有助于逻辑与序列推理任务。这些效应在认知负荷标注、表面文本特征及编辑模式分析中均具鲁棒性,并能跨优化框架泛化。整体表明,提示优化失败源于编辑类型与任务特征之间的系统性交互,而非随机优化产物,为优化器行为提供了特征级刻画,推动未来任务条件化优化器设计。

原文摘要 · Abstract (English)

Automated prompt optimization methods (e.g., DSpy, TextGrad) can substantially improve the performance of large language model (LLM), however, their generalization ability across different tasks remains underperformed. In practice, the superiority of the optimized prompt on one benchmark often fails to transfer to another, and this limitation persists even when switching across different LLM backbones. To investigate the underexplored sources of heterogeneity in prompt performance, we conduct a causal inference-inspired observational analysis of optimized prompts across a diverse set of optimization frameworks, LLM backbones, and NLP benchmarks. To achieve the goal, we build upon the propensity-adjusted associational analysis together with multiple complementary representations of prompt edits, where the consistent task-conditioned edits patterns are identified. We find that complexity-increasing and meta-instructional edits are negatively associated with mathematical and multi-hop reasoning performance, whereas step-by-step and meta-cognitive edits improve logical and sequential reasoning tasks. These effects are robust across cognitive-load annotations, surface-level text features, and edit-motif analyses, and can generalize across optimization frameworks. Overall, these results indicate that prompt optimization failures arise from systematic interactions between edit families and task characteristics rather than random optimization artifacts, providing feature-level characterization of optimizer behavior and motivating future task-conditioned optimizer design.

提示优化大模型因果分析任务适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。