arXiv:2505.10736cs.CL2025-05ACL被引 6

用模型实时性能指导数据选择,提升提示词优化效果与稳定性。

Model Performance-Guided Evaluation Data Selection for Effective Prompt Optimization

  • 基于语义聚类和边界分析选样,再用模型表现迭代替换冗余样本。
  • 在BIG-bench上提升效果1.6%~5.3%,稳定性提高至少57%。
  • 适用于新数据集,计算开销低于1%,可增强现有选子集方法。

优化大语言模型性能需精心设计提示词,但手动工程耗时且常无效。自动化提示优化技术虽能缓解此问题,但多数依赖随机选取的评估子集,难以代表全数据集,导致评估不可靠、提示词优化效果不佳。现有核集(coreset)选择方法因聚类困难、数据收集成本高及缺乏新或私有数据集的性能数据,不适用于提示优化。为此,我们提出IPOMP:一种基于实时模型性能的迭代评估数据选择方法,用于高效提示优化。该方法分两阶段:先通过语义聚类和边界分析选取代表性与多样性样本,再结合实时模型性能数据迭代优化,替换冗余样本。在BIG-bench上的实验表明,相比最先进基线,IPOMP提升效果1.6%至5.3%,稳定性提升至少57%,计算开销低于1%。结果还显示,该实时性能引导的精炼策略可普遍应用于增强现有核集选择方法。

原文摘要 · Abstract (English)

Optimizing Large Language Model (LLM) performance requires well-crafted prompts, but manual prompt engineering is labor-intensive and often ineffective. Automated prompt optimization techniques address this challenge but the majority of them rely on randomly selected evaluation subsets, which fail to represent the full dataset, leading to unreliable evaluations and suboptimal prompts. Existing coreset selection methods, designed for LLM benchmarking, are unsuitable for prompt optimization due to challenges in clustering similar samples, high data collection costs, and the unavailability of performance data for new or private datasets. To overcome these issues, we propose IPOMP, an Iterative evaluation data selection for effective Prompt Optimization using real-time Model Performance. IPOMP is a two-stage approach that selects representative and diverse samples using semantic clustering and boundary analysis, followed by iterative refinement with real-time model performance data to replace redundant samples. Evaluations on the BIG-bench dataset show that IPOMP improves effectiveness by 1.6% to 5.3% and stability by at least 57% compared with SOTA baselines, with minimal computational overhead below 1%. Furthermore, the results demonstrate that our real-time performance-guided refinement approach can be universally applied to enhance existing coreset selection methods.

提示优化模型评估数据选择LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。