用户与大模型对话时语言风格不同,需调整训练策略以提升交互体验。
Mind the Gap: Linguistic Divergence and Adaptation Strategies in Human-LLM Assistant vs. Human-Human Interactions
- 对比人类与大模型交互,发现语法流畅度、礼貌程度和词汇多样性差异显著。
- 使用风格多样的数据训练的模型表现更优,优于仅用原始或单一风格数据训练的模型。
- 部署后通过语句重构提升效果有限,建议在训练阶段引入多样风格数据。
随着大型语言模型(LLMs)在客户服务场景中的广泛应用,一个关键但未被充分研究的问题是:用户与聊天机器人互动时的沟通方式是否不同于与真人客服的交流。本研究提供了实证证据,表明用户在与聊天机器人和真人客服互动时采用不同的沟通风格。分析显示,两种情境下用户的语法流畅度、礼貌程度及词汇多样性存在显著差异。这说明仅基于人-人交互数据训练的模型可能无法适应部署后出现的沟通风格变化。为增强模型对上线后沟通风格改变的鲁棒性,我们尝试了两种策略:(1) 在后训练阶段进行数据增强;(2) 推理时对用户输入进行重写。结果表明,使用风格多样数据训练的模型显著优于仅使用原始或风格一致数据训练的模型,而推理时重写的效果较弱。这些发现有助于优化模型以改善人-大模型交互体验。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) are increasingly deployed in customer-facing applications, a critical yet underexplored question is how users communicate differently with LLM chatbots compared to human agent. In this study, we present empirical evidence that users adopt distinct communication styles when users interact with chatbots versus human agents. Our analysis reveals significant differences in grammatical fluency, politeness, and lexical diversity in user language between the two settings. These findings suggest that models trained exclusively on human-human interaction data may not adequately accommodate the communication style shift that occurs once an LLM chatbot is deployed. To enhance LLM robustness to post-launch communication style changes, we experimented with two strategies: (1) data augmentation during the post-training phase and (2) inference-time user message reformulation. Our results indicate that models trained on stylistically diverse datasets significantly outperform those trained exclusively on original or stylistically uniform datasets, while inference-time reformulation proved less effective. These insights help us to better adapt our models for improved LLM-user interaction experiences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。