arXiv:2603.25169cs.CL2026-03

自动提示优化能否替代专家设计?实验发现效果接近,视任务而定。

To Write or to Automate Linguistic Prompts, That Is the Question

  • 对比人工、基础和优化后的提示模板在三类语言任务中的表现
  • 术语插入任务中,自动与人工提示效果无显著差异
  • 专家提示更擅于发现错误,自动优化更擅长描述问题特征

大语言模型性能高度依赖提示设计,但自动提示优化是否可替代专家提示工程仍未知。本文首次系统比较了人工零样本专家提示、基础DSPy签名以及GEPA优化的DSPy签名在翻译、术语插入和语言质量评估任务中的表现,涵盖五种模型配置。结果因任务而异:术语插入任务中,优化与人工提示效果统计上无差异;翻译任务中,各方法在不同模型上各有优势;语言质量评估中,专家提示更强于错误检测,而优化方法更优在问题表征。整体上,GEPA能显著提升初始最小化DSPy签名,多数专家优化对比未见显著差异。需注意:该比较存在不对称性——GEPA基于标准标注集搜索,而专家提示无需标签数据,仅依赖领域知识与迭代打磨。

原文摘要 · Abstract (English)

LLM performance is highly sensitive to prompt design, yet whether automatic prompt optimization can replace expert prompt engineering in linguistic tasks remains unexplored. We present the first systematic comparison of hand-crafted zero-shot expert prompts, base DSPy signatures, and GEPA-optimized DSPy signatures across translation, terminology insertion, and language quality assessment, evaluating five model configurations. Results are task-dependent. In terminology insertion, optimized and manual prompts produce mostly statistically indistinguishable quality. In translation, each approach wins on different models. In LQA, expert prompts achieve stronger error detection while optimization improves characterization. Across all tasks, GEPA elevates minimal DSPy signatures, and the majority of expert-optimized comparisons show no statistically significant difference. We note that the comparison is asymmetric: GEPA optimization searches programmatically over gold-standard splits, whereas expert prompts require in principle no labeled data, relying instead on domain expertise and iterative refinement.

提示工程自动优化语言任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。