arXiv:2608.17605cs.CLcs.AI2026-08被引 2

对话系统正从单轮文本转向多轮跨模态交互,需持续记忆与上下文理解。

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

  • 构建多轮对话的跨模态记忆与上下文保持机制
  • 当前系统在跨轮次对齐与全双工交互上仍存在明显短板
  • 适合研究多轮对话、跨模态融合与智能体交互的学者参考

对话人工智能正从孤立的文本输入演进为持续的多模态交互。真实对话中,用户会修正目标、中断回复、切换话题或引入新信息,要求系统能维持跨轮次上下文。这使多轮对话成为独特挑战,需具备持续记忆、跨模态/工具/外部知识对齐及跨语言文化适应能力。本文综述了纯文本对话、AudioLLM与语音原生系统、多模态与全模态系统、以及工具增强型智能体的研究进展,围绕数据集与基准、建模范式、训练策略、评估设置和共性挑战进行组织分析。结果显示,多模态感知与生成能力发展快于持续对话连贯性支持。尽管系统在跨模态表达与行动上更强,但在持久记忆、跨轮对齐、全双工交互、鲁棒评估与文化适配方面仍存不足。最后提出一个研究议程,旨在实现能跨轮次、跨模态、跨文化记住、修正、对齐、说话、倾听、行动与适应的对话系统。

原文摘要 · Abstract (English)

Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context across turns. This makes multi-turn dialogue a distinct challenge requiring systems to maintain and update memory, ground responses across modalities, tools, and external knowledge, and adapt across languages and cultures. This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents. We organize the literature around datasets and benchmarks, modeling paradigms, training strategies, evaluation setups, and cross-cutting challenges. Our analysis shows that support for multiple modalities has advanced faster than the ability to sustain coherent interaction across a session. Despite stronger capabilities to perceive, speak, and act across modalities, current systems still struggle with persistent memory, cross-turn grounding, full-duplex interaction, robust evaluation, and cultural alignment. We conclude with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures. (https://github.com/faiza-sfa/multiturn-conversational-ai-survey)

多轮对话跨模态智能体评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。