arXiv:2409.08330cs.CLcs.CY2024-09被引 19

用10万组对话对比人与大模型的对话差异,发现模拟效果不理想。

Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Dialogue

  • 构建10万对真实与大模型生成的对话数据集
  • 发现大模型生成内容与真人对话匹配度较低
  • 适合研究对话真实性或评估生成模型的学者

由于需招募、培训并收集参与者数据,对话任务的研究与数据集构建成本高昂。为此,近期许多工作尝试使用大语言模型(LLMs)模拟人类-人类及人类-大模型互动,因其在多种场景下能生成看似自然的人类语言。然而,这些模拟在多大程度上真正反映真实对话?本研究基于WildChat数据集,生成10万对大模型-大模型与人类-大模型对话,并量化其与真人对话的对齐程度。结果表明,模拟对话与真人对话在风格、内容等多个文本属性上存在系统性偏差,整体对齐度较低。此外,在英语、中文和俄语对话中,模型表现相近。结果显示,当真人自身表达风格更接近大模型时,模拟效果更好。

原文摘要 · Abstract (English)

Studying and building datasets for dialogue tasks is both expensive and time-consuming due to the need to recruit, train, and collect data from study participants. In response, much recent work has sought to use large language models (LLMs) to simulate both human-human and human-LLM interactions, as they have been shown to generate convincingly human-like text in many settings. However, to what extent do LLM-based simulations \textit{actually} reflect human dialogues? In this work, we answer this question by generating a large-scale dataset of 100,000 paired LLM-LLM and human-LLM dialogues from the WildChat dataset and quantifying how well the LLM simulations align with their human counterparts. Overall, we find relatively low alignment between simulations and human interactions, demonstrating a systematic divergence along the multiple textual properties, including style and content. Further, in comparisons of English, Chinese, and Russian dialogues, we find that models perform similarly. Our results suggest that LLMs generally perform better when the human themself writes in a way that is more similar to the LLM's own style.

对话模拟大模型评估真实性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。