arXiv:2409.18602cs.CL2024-09EMNLP被引 13

诊断大模型在多人对话中识别回应对象与选择回复的弱点

Do LLMs suffer from Multi-Party Hangover? A Diagnostic Approach to Addressee Recognition and Response Selection in Conversations

  • 构建结构多样化的子数据集,按对话复杂度分析模型表现
  • 发现回复选择依赖文本内容,而对象识别需把握对话结构
  • 揭示不同任务对提示敏感度差异,适合对话系统研究者参考

评估多人对话(MPC)系统性能面临挑战,因语言特征与对话结构紧密关联。传统评估方法常忽略模型在不同结构复杂度下的行为差异。本文提出方法论流程,针对响应选择与对话对象识别任务进行诊断分析。基于大规模公开在线多人对话语料,提取具有固定用户数和丰富结构多样性的诊断子数据集。工作遵循数据最小化原则,避免使用原始用户名以保护隐私,并提出替代原始消息文本的方法。结果表明,响应选择更依赖对话文本内容,而对象识别需捕捉其结构维度。在零样本设置下使用大模型进一步显示,任务对提示变化的敏感性存在差异。

原文摘要 · Abstract (English)

Assessing the performance of systems to classify Multi-Party Conversations (MPC) is challenging due to the interconnection between linguistic and structural characteristics of conversations. Conventional evaluation methods often overlook variances in model behavior across different levels of structural complexity on interaction graphs. In this work, we propose a methodological pipeline to investigate model performance across specific structural attributes of conversations. As a proof of concept we focus on Response Selection and Addressee Recognition tasks, to diagnose model weaknesses. To this end, we extract representative diagnostic subdatasets with a fixed number of users and a good structural variety from a large and open corpus of online MPCs. We further frame our work in terms of data minimization, avoiding the use of original usernames to preserve privacy, and propose alternatives to using original text messages. Results show that response selection relies more on the textual content of conversations, while addressee recognition requires capturing their structural dimension. Using an LLM in a zero-shot setting, we further highlight how sensitivity to prompt variations is task-dependent.

对话系统大模型诊断多角色对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。