用多智能体生成高质量对话数据,提升AI对话系统训练效果。
ConvoGen: Enhancing Conversational AI with Synthetic Data: A Multi-Agent Approach
- 通过动态更新的少量样本池进行迭代采样,生成多样对话
- 生成数据在意图分类等任务中表现优异,有效增强模型性能
- 适合需要大量真实对话数据的研究者与开发者使用
本文提出ConvoGen:一种基于多智能体系统的合成对话数据生成框架。该方法利用少样本学习,并通过动态更新的少样本中心进行迭代采样,生成多样化且真实的对话场景。生成的数据可用于训练和评估对话AI模型,也可用于扩充现有数据集,支持对话意图识别、对话摘要等任务。实验表明,该方法能有效生成高质量、多样化的合成对话数据,显著提升对话AI系统的开发与评估能力。
原文摘要 · Abstract (English)
In this paper, we present ConvoGen: an innovative framework for generating synthetic conversational data using multi-agent systems. Our method leverages few-shot learning and introduces iterative sampling from a dynamically updated few-shot hub to create diverse and realistic conversational scenarios. The generated data has numerous applications, including training and evaluating conversational AI models, and augmenting existing datasets for tasks like conversational intent classification or conversation summarization. Our experiments demonstrate the effectiveness of this method in producing high-quality diverse synthetic conversational data, highlighting its potential to enhance the development and evaluation of conversational AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。