arXiv:2506.10504cs.CLcs.AI2025-06EMNLP被引 3

测试大模型在多人对话中的状态追踪能力,发现性能显著下降。

Beyond Single-User Dialogue: Assessing Multi-User Dialogue State Tracking Capabilities of Large Language Models

  • 基于语用理论生成第二名用户话语,构建多用户对话数据
  • 大模型在多人场景下表现远低于单人场景,状态追踪能力明显减弱
  • 为真实复杂对话研究提供新评估基准,适合对话系统开发者

大语言模型(LLMs)在零样本对话状态追踪(DST)中表现出色,减少了对特定任务训练的需求。然而,传统DST基准主要关注结构化的用户-代理对话,未能反映真实世界中多人互动的复杂性。本文在降低数据构建成本的前提下,评估了LLMs在多用户DST中的鲁棒性。受近期基于LLM的数据标注进展启发,我们通过语用理论扩展现有DST数据集,生成第二名用户的发言。该方法系统地将第二名用户的对话融入原有会话中,实现对多用户场景下LLMs的可控评估。实验结果表明,与单用户DST相比,模型性能显著下降,凸显当前LLMs在多说话人环境中提取和追踪对话状态的局限性。研究强调需推动未来工作提升模型在多用户场景下的表现,为更真实、更鲁棒的DST模型铺平道路。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable performance in zero-shot dialogue state tracking (DST), reducing the need for task-specific training. However, conventional DST benchmarks primarily focus on structured user-agent conversations, failing to capture the complexities of real-world multi-user interactions. In this study, we assess the robustness of LLMs in multi-user DST while minimizing dataset construction costs. Inspired by recent advances in LLM-based data annotation, we extend an existing DST dataset by generating utterances of a second user based on speech act theory. Our methodology systematically incorporates a second user's utterances into conversations, enabling a controlled evaluation of LLMs in multi-user settings. Experimental results reveal a significant performance drop compared to single-user DST, highlighting the limitations of current LLMs in extracting and tracking dialogue states amidst multiple speakers. Our findings emphasize the need for future research to enhance LLMs for multi-user DST scenarios, paving the way for more realistic and robust DST models.

对话状态追踪大模型多用户对话评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。