arXiv:2603.17094cs.CL2026-03

评估大模型对话中不一致与对抗行为,揭示其难以真实模拟人类社交复杂性。

Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction

  • 用大模型做裁判,逐轮检测10类对话异常行为。
  • 原生提示下,模型对话中的冲突行为远少于真人。
  • 调参无法稳定控制行为,微调反而导致重复等单一问题。

使用大语言模型(LLMs)模拟人类对话已成为建模人际互动的可扩展方法。然而,由于人类对话天然包含不一致和不协作行为(如误解、打断),模拟此类行为极具挑战。目前对人类与模型生成对话中不一致与不协作行为的对比分析仍有限,而重现这些行为对于构建拟人化复杂社交互动至关重要。本文提出 CoCoEval 框架,利用大模型作为评判者,在话轮层面检测10类不一致与不协作行为。我们评估 GPT-4.1、GPT-5.1 与 Claude Opus 4,对比其在学术、商业、政府会议及辩论场景中生成对话与真人对话的行为频率。结果表明:(1)在原始提示下,模型对话中不一致与不协作行为显著少于真人;(2)提示工程无法可靠控制这些行为,不同提示导致行为过度或不足;(3)在真人对话上进行监督微调后,模型反而过度产生如重复等少数特定行为。研究揭示了模拟人类对话的困难,警示大模型不宜作为人类社交互动的代理。

原文摘要 · Abstract (English)

Simulating human conversations using large language models (LLMs) has emerged as a scalable methodology for modeling human social interaction. However, simulating human conversations is challenging because they inherently involve inconsistent and uncollaborative behaviors, such as misunderstandings and interruptions. Analysis comparing inconsistent and uncollaborative behaviors in human- and LLM-generated conversations remains limited, although reproducing these behaviors is integral to simulating human-like and complex social interaction. In this work, we introduce CoCoEval, an evaluation framework that analyzes LLM-simulated conversations by detecting 10 types of inconsistent and uncollaborative behaviors at the turn level using an LLM-as-a-Judge. Using CoCoEval, we evaluate GPT-4.1, GPT-5.1, and Claude Opus 4 by comparing the frequencies of detected behaviors in conversations simulated by each model and in human conversations across academic, business, and governmental meetings, as well as debates. Our analysis shows that (1) under vanilla prompting, LLM-simulated conversations exhibit far fewer inconsistent and uncollaborative behaviors than human conversations; (2) prompt engineering does not provide reliable control over these behaviors, as our results show that different prompts lead to their under- or overproduction; and (3) supervised fine-tuning on human conversations can lead LLMs to overproduce a narrow set of behaviors, such as repetition. Our findings highlight the difficulty of simulating human conversations, raising concerns about the use of LLMs as a proxy for human social interaction.

大模型评估对话模拟社会行为行为检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。