arXiv:2506.21138cs.SEcs.AI2025-06被引 3

用多样本提示和强化学习优化提示,生成更多样且实用的合成数据。

Multi-Sample Prompting and Actor-Critic Prompt Optimization for Diverse Synthetic Data Generation

  • 通过多样本提示减少生成内容重复性,提升多样性。
  • 在4个需求工程任务中,F1分数最高提升43.8个百分点。
  • 合成数据可媲美甚至超越人工标注数据,尤其适合数据稀缺场景。

高质量标注数据是训练和评估机器学习模型的基础,但医疗与需求工程(RE)等领域受限于数据稀缺、隐私保护或专有权限。尽管大语言模型(LLMs)为合成数据生成(SDG)提供了前景,但其生成内容往往重复且多样性不足,影响下游任务效果。本文提出两种改进策略:(1) 多样本提示,即每条提示生成多个样本以降低重复;(2) 带有演员-评论家机制的提示优化(PACE),通过迭代优化提示以最大化多样性。将二者集成至基于特征模型的可配置合成数据生成器Synthline,评估其在四个RE分类任务中的表现。多样本提示显著提升多样性和下游性能,F1分数提升6至43.8个百分点。基于PACE的提示优化虽提升词汇多样性,但对任务性能影响因任务而异,揭示了单纯优化多样性可能带来的风险。最显著的是,在真实标注数据有限的任务中,合成数据可达到甚至超过人工数据表现,F1分数最高提升15.4个百分点。

原文摘要 · Abstract (English)

High-quality labeled datasets are fundamental for training and evaluating machine learning models, yet domains such as healthcare and Requirements Engineering (RE) face persistent barriers due to data scarcity, privacy constraints, or proprietary restrictions. While Large Language Models (LLMs) offer a promising avenue for Synthetic Data Generation (SDG), LLM-generated data tends to be repetitive and low in diversity, reducing its effectiveness for downstream tasks. Two approaches show potential for addressing this limitation: (1) multi-sample prompting, which generates multiple samples per prompt to reduce repetition, and (2) Prompt with Actor-Critic Editing (PACE), which iteratively refines prompts to maximize diversity. We integrate both mechanisms into Synthline, a Feature Model-based configurable synthetic data generator, and assess their effects on diversity and downstream utility across four RE classification tasks. Multi-sample prompting consistently improves both diversity and utility, with F1-score gains of 6 to 43.8 percentage points. PACE-based prompt optimization consistently improves lexical diversity but produces task-dependent utility effects, revealing the risks of optimizing for diversity alone. Most notably, synthetic data can match or surpass human-authored data for tasks where real labeled data is limited, with improvements of up to 15.4 percentage points in F1-score.

合成数据大模型提示优化需求工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。