arXiv:2509.17619cs.IR2025-09被引 9

对比真人与大模型模拟用户在任务对话中的表现差异。

Human vs. Agent in Task-Oriented Conversations

  • 构建三维度十指标框架,系统分析对话策略与风格
  • 发现两类用户在提问广度、反馈倾向等多方面存在显著差异
  • 为大模型用户仿真提供可复用的行为分析框架,适合对话系统研究者

任务导向型对话系统对高效满足用户需求至关重要,但其开发依赖大量高质量对话数据,获取成本高。尽管大语言模型(LLM)在生成合成对话方面展现潜力,但其生成的代理用户能否有效替代真实人类用户仍不明确。本文首次系统比较了LLM模拟用户与真实人类用户在个性化任务对话中的表现。我们提出涵盖对话策略、交互风格和评估三个维度的综合分析框架,包含十个具体维度,并在四个典型场景下采集了平行对话数据集,确保条件一致。分析显示,两类用户在问题解决方式、提问范围、用户参与度、上下文依赖性、反馈极性与承诺度、语言风格及幻觉意识等方面存在显著差异。但在深度优先或广度优先策略以及有用性维度上表现出一致性。这些发现为推进基于LLM的用户仿真提供了关键洞见。所构建的多维分类体系形成通用化用户行为分析框架,既揭示了代理用户与人类用户的行为模式,也为未来对话系统中用户仿真的优化提供了新视角。

原文摘要 · Abstract (English)

Task-oriented conversational systems are essential for efficiently addressing diverse user needs, yet their development requires substantial amounts of high-quality conversational data that is challenging and costly to obtain. While large language models (LLMs) have demonstrated potential in generating synthetic conversations, the extent to which these agent-generated interactions can effectively substitute real human conversations remains unclear. This work presents the first systematic comparison between LLM-simulated users and human users in personalized task-oriented conversations. We propose a comprehensive analytical framework encompassing three key aspects (conversation strategy, interaction style, and conversation evaluation) and ten distinct dimensions for evaluating user behaviors, and collect parallel conversational datasets from both human users and LLM agent users across four representative scenarios under identical conditions. Our analysis reveals significant behavioral differences between the two user types in problem-solving approaches, question broadness, user engagement, context dependency, feedback polarity and promise, language style, and hallucination awareness. We found consistency in the agent users and human users across the depth-first or breadth-first dimensions, as well as the usefulness dimensions. These findings provide critical insights for advancing LLM-based user simulation. Our multi-dimensional taxonomy constructed a generalizable framework for analyzing user behavior patterns, offering insights from LLM agent users and human users. By this work, we provide perspectives on rethinking how to use user simulation in conversational systems in the future.

对话系统用户仿真大模型行为分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。