用困惑度先验让小数据下提示词优化更泛化,理论更紧致。
Prompts Generalize with Low Data: Non-vacuous Generalization Bounds for Optimizing Prompts with More Informative Priors
- 引入困惑度作为分布相关先验,约束提示词搜索空间。
- 在少量数据下仍得到非平凡泛化界,优于传统方法。
- 适合低资源场景的提示工程研究者与实践者。
许多提示工程方法在实践中表现良好,即使在任务数据极少的情况下优化大规模提示空间亦然。近期工作通过将PAC-Bayes理论应用于离散提示空间,部分解释了这一成功,但其泛化界仅在数据丰富时才非平凡。我们主张,这种广泛成功可通过更仔细地考虑数据或分布相关的困惑度来更好解释,该困惑度充当有效先验,引导优化趋向于更符合任务的“自然”提示。本文推导出新的泛化界,在数据稀缺时依然非平凡,形式化分析了困惑度正则化如何通过限制探索来收紧边界。实验上,我们验证了这些边界的有效性及困惑度正则化的实际收益,显著提升了提示词的泛化性能。
原文摘要 · Abstract (English)
Many prompt engineering techniques have been successful in practice, even when optimizing over a large prompt space with with a small amount of task-specific data. Recent work has partially explained this success by showing generalization bounds which apply PAC-Bayes theory to the discrete prompt space, but they are non-vacuous only in data-rich scenarios. We argue that such widespread success can be more fully explained through more carefully considering data- or distribution-dependent perplexity, which acts as an effective prior and steers the optimization towards prompts that are more ``natural'' for the task at hand. We derive novel generalization bounds that are non-vacuous for data-scarce prompt optimization via more useful priors, formally analyzing how perplexity regularization tightens these bounds by limiting exploration. Empirically, we explore both the bounds' effectiveness and the practical benefits of perplexity regularization in improving prompt generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。