arXiv:2509.05375cs.AI2025-09

揭示提示工程优化景观的复杂结构差异

Characterizing Fitness Landscape Structures in Prompt Engineering

  • 通过语义嵌入空间自相关分析,对比两种提示生成策略的景观拓扑
  • 系统枚举产生平滑衰减相关性,多样化生成呈现中间距离峰值的崎岖结构
  • 发现不同错误类型对应的景观崎岖程度不同,为优化算法设计提供依据

尽管提示工程已成为提升大模型性能的关键技术,其背后的优化景观仍不明确。现有方法将提示优化视为黑箱问题,应用复杂搜索算法却未刻画所处景观的拓扑结构。本文通过在语义嵌入空间中进行自相关分析,系统研究提示工程中的适应度景观结构。在两类提示生成策略(系统枚举:1,024 个提示;新颖性驱动多样性:1,000 个提示)的误差检测任务上实验发现:系统枚举生成的景观呈现平滑衰减的自相关,而多样化生成则表现出非单调模式,在中等语义距离处出现相关性峰值,表明存在崎岖且分层的景观结构。对 10 种不同错误类型的任务特异性分析进一步揭示了各类错误对应景观的崎岖程度差异。本研究为理解提示工程优化的复杂性提供了实证基础。

原文摘要 · Abstract (English)

While prompt engineering has emerged as a crucial technique for optimizing large language model performance, the underlying optimization landscape remains poorly understood. Current approaches treat prompt optimization as a black-box problem, applying sophisticated search algorithms without characterizing the landscape topology they navigate. We present a systematic analysis of fitness landscape structures in prompt engineering using autocorrelation analysis across semantic embedding spaces. Through experiments on error detection tasks with two distinct prompt generation strategies -- systematic enumeration (1,024 prompts) and novelty-driven diversification (1,000 prompts) -- we reveal fundamentally different landscape topologies. Systematic prompt generation yields smoothly decaying autocorrelation, while diversified generation exhibits non-monotonic patterns with peak correlation at intermediate semantic distances, indicating rugged, hierarchically structured landscapes. Task-specific analysis across 10 error detection categories reveals varying degrees of ruggedness across different error types. Our findings provide an empirical foundation for understanding the complexity of optimization in prompt engineering landscapes.

提示工程优化景观语义嵌入自相关分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。