综述多角色对话智能体如何理解心理状态、语义与对话流
Multi-Party Conversational Agents: A Survey
- 从心理状态建模、语义理解到对话行为预测,系统梳理三大核心挑战
- 指出心智理论(ToM)是构建智能多角色对话系统的关键
- 适合对多智能体交互、大模型对话应用感兴趣的科研与工程人员
多角色对话智能体(MPCAs)是能够同时与超过两人进行对话的系统。与传统双人对话系统不同,MPCAs的设计需兼顾话语语义和社交动态的理解。本文通过三个核心问题综述近年来的研究进展:1)能否建模参与者的心理状态?(心理状态建模);2)能否准确理解对话内容?(语义理解);3)能否推理并预测未来的对话发展?(智能体行为建模)。文章回顾了从经典机器学习到大语言模型(LLMs)及多模态系统的方法体系。分析表明,心智理论(ToM)对于构建高智能的MPCAs至关重要,而多模态理解虽具潜力但仍处于探索阶段。最后,本文为未来研究者提供了发展方向建议。
原文摘要 · Abstract (English)
Multi-party Conversational Agents (MPCAs) are systems designed to engage in dialogue with more than two participants simultaneously. Unlike traditional two-party agents, designing MPCAs faces additional challenges due to the need to interpret both utterance semantics and social dynamics. This survey explores recent progress in MPCAs by addressing three key questions: 1) Can agents model each participants' mental states? (State of Mind Modeling); 2) Can they properly understand the dialogue content? (Semantic Understanding); and 3) Can they reason about and predict future conversation flow? (Agent Action Modeling). We review methods ranging from classical machine learning to Large Language Models (LLMs) and multi-modal systems. Our analysis underscores Theory of Mind (ToM) as essential for building intelligent MPCAs and highlights multi-modal understanding as a promising yet underexplored direction. Finally, this survey offers guidance to future researchers on developing more capable MPCAs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。