用大模型生成多样化虚拟用户,模拟真实对话场景。
Simulating User Diversity in Task-Oriented Dialogue Systems using Large Language Models
- 用大模型构建包含多元背景和目标的虚拟用户画像
- GPT-o1生成的用户分布更均匀,GPT-4o则更偏向特定类型
- 适用于测试对话系统在复杂用户场景下的表现
本研究探索利用大语言模型(LLMs)生成合成用户,并模拟用户与任务导向对话系统的交互过程,提供详细结果与分析。提出一种新型用户仿真方法,通过大模型创建多样化的用户资料,包括不同人口统计特征、多重对话目标、各异的交流风格、初始知识水平、兴趣点及对话目的。采用两种专有大模型GPT-4o和GPT-o1(Achiam et al., 2023)生成用户群体,该群体在多维度上表现出显著异质性。对生成的用户画像进行深入分析,评估其多样性、一致性及潜在偏见。结果显示,GPT-o1在多数用户属性上生成更广泛且均衡的分布,而GPT-4o则表现出更强的分布偏斜。这些生成的用户资料随后被用于与任务导向对话系统交互,模拟多轮对话会话。
原文摘要 · Abstract (English)
In this study, we explore the application of Large Language Models (LLMs) for generating synthetic users and simulating user conversations with a task-oriented dialogue system and present detailed results and their analysis. We propose a comprehensive novel approach to user simulation technique that uses LLMs to create diverse user profiles, set goals, engage in multi-turn dialogues, and evaluate the conversation success. We employ two proprietary LLMs, namely GPT-4o and GPT-o1 (Achiam et al., 2023), to generate a heterogeneous base of user profiles, characterized by varied demographics, multiple user goals, different conversational styles, initial knowledge levels, interests, and conversational objectives. We perform a detailed analysis of the user profiles generated by LLMs to assess the diversity, consistency, and potential biases inherent in these LLM-generated user simulations. We find that GPT-o1 generates more heterogeneous user distribution across most user attributes, while GPT-4o generates more skewed user attributes. The generated set of user profiles are then utilized to simulate dialogue sessions by interacting with a task-oriented dialogue system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。