arXiv:2504.04717cs.CLcs.AI2025-04综述被引 80

系统梳理大模型多轮对话的评估与优化方法,覆盖从数学编程到医疗教育等场景。

Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models

  • 按任务类型分类,梳理多轮对话挑战与评估框架。
  • 总结模型内优化、外部记忆增强与智能体协作等多种提升策略。
  • 适合关注对话系统鲁棒性与实际应用的研究者参考。

大语言模型在单轮任务上表现显著提升,但现实应用对复杂多轮交互需求日益增长。本文系统综述了近期在评估与增强多轮大模型交互方面的进展。基于任务导向的分类体系,涵盖数学与编程中的指令遵循,以及角色扮演、医疗、教育和对抗性越狱等场景中的对话参与度。深入分析长期对话中保持上下文连贯性、公平性与响应性的挑战。将现有基准与数据集归类为反映多轮对话评估演进的清晰类别,并综述多种增强方法:包括模型内部策略(上下文学习、监督微调、强化学习与架构创新)、外部集成方法(记忆增强、检索机制、知识图谱)以及用于协作交互的智能体技术。最后,指出当前开放问题与未来研究方向,以进一步提升多轮交互的鲁棒性与有效性。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have substantially improved single-turn task performance, yet real-world applications increasingly demand sophisticated multi-turn interactions. This survey provides a comprehensive review of recent progress in evaluating and enhancing multi-turn LLM interactions. Centered on a task-oriented taxonomy-spanning instruction following in domains such as mathematics and coding, and conversational engagement in role-playing, healthcare, education, and adversarial jailbreak settings-we systematically examine the challenges of maintaining context, coherence, fairness, and responsiveness across prolonged dialogues. We organize existing benchmarks and datasets into coherent categories reflecting the evolving landscape of multi-turn dialogue evaluation, and review a broad spectrum of enhancement methodologies, including model-centric strategies (in-context learning, supervised fine-tuning, reinforcement learning, and architectural innovations), external integration approaches (memory augmentation, retrieval-based methods, and knowledge graphs), and agent-based techniques for collaborative interaction. Finally, we identify open challenges and promising directions for future research to further improve the robustness and effectiveness of multi-turn LLM interactions.

多轮对话大模型评估智能体协作任务导向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。