arXiv:2412.18196cs.CLcs.LG2024-12被引 3

自动提示生成易受干扰,新方法用伪梯度提升鲁棒性

Auto-Prompt Generation is Not Robust: Prompt Optimization Driven by Pseudo Gradient

  • 用扰动类型做伪梯度信号,无须真实梯度优化提示
  • 在多种模型和任务上,扰动下性能优于现有方法
  • 首次系统评测自动提示的鲁棒性,适合关注模型稳定性的研究者

尽管自动提示生成方法受到广泛关注,其鲁棒性仍不明确。本文提出PertBench,一个包含多种输入扰动的综合性基准数据集,用于系统评估现有自动提示技术的鲁棒性。分析发现,现有提示生成策略存在显著脆弱性:即使提示发生微小变化,模型输出也会出现显著差异。为解决此问题,我们提出PGO,一种基于扰动类型的伪梯度信号引导的无梯度提示生成框架,使大模型生成更具鲁棒性的提示。与仅在干净、结构化输入上评估质量的方法不同,我们的方法显式强调在噪声和扰动条件下的表现。在多个任务和多种大模型上的大量实验表明,PGO在输入扰动下始终优于先前方法。

原文摘要 · Abstract (English)

While automatic prompt generation methods have recently received significant attention, their robustness remains poorly understood. In this paper, we introduce PertBench, a comprehensive benchmark dataset that includes a wide range of input perturbations, designed to systematically evaluate the robustness of current auto-prompting techniques. Our analysis reveals substantial vulnerabilities in existing prompt generation strategies, where even minor modifications to the prompt can lead to significant differences in model output. To address this issue, we propose PGO, a gradient-free prompt generation framework that leverages perturbation types as pseudo-gradient signals to guide LLMs in producing more robust prompts. In contrast to existing methods that assess prompt quality only on clean, well-structured inputs, our approach explicitly emphasizes robustness under noisy and perturbed conditions. Extensive experiments across diverse tasks and multiple LLMs show PGO consistently outperforms previous methods in maintaining performance under input perturbations.

提示工程鲁棒性大模型无梯度优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。