arXiv:2501.16643cs.CLcs.AI2025-01中稿 · presentation at In…被引 15

评测大模型在三人对话中识别说话对象的能力,发现效果仅略高于随机。

An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue

  • 构建三人群聊多模态语料库,标注20%对话回合的明确发言对象。
  • 用GPT-4o测试发现其识别准确率仅略高于随机水平。
  • 揭示大模型在多人对话中理解指向关系的严重不足,适合对话系统研究者参考。

处理多人对话是推进语音对话系统的重要一步,需要针对多人互动设计特定任务。为此,我们构建了一个包含三人讨论的多模态多人对话语料库。本文聚焦于发言对象识别任务,即判断下一轮发言应针对谁,这是多人对话系统中的关键组件。语料库的一个子集被标注了发言对象信息,结果显示约20%的对话回合存在明确发言对象。为评估任务复杂性,我们对大型语言模型GPT-4o进行了基准测试,结果表明其准确率仅略高于随机水平,凸显了在多人对话中进行发言对象识别的挑战。这些发现强调了需进一步研究以提升大模型对多人对话动态的理解与导航能力。

原文摘要 · Abstract (English)

Handling multi-party dialogues represents a significant step for advancing spoken dialogue systems, necessitating the development of tasks specific to multi-party interactions. To address this challenge, we are constructing a multi-modal multi-party dialogue corpus of triadic (three-participant) discussions. This paper focuses on the task of addressee recognition, identifying who is being addressed to take the next turn, a critical component unique to multi-party dialogue systems. A subset of the corpus was annotated with addressee information, revealing that explicit addressees are indicated in approximately 20% of conversational turns. To evaluate the task's complexity, we benchmarked the performance of a large language model (GPT-4o) on addressee recognition. The results showed that GPT-4o achieved an accuracy only marginally above chance, underscoring the challenges of addressee recognition in multi-party dialogue. These findings highlight the need for further research to enhance the capabilities of large language models in understanding and navigating the intricacies of multi-party conversational dynamics.

对话系统多模态大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。