arXiv:2512.02593cs.CLcs.MA2025-12EMNLP被引 5

将大语言模型用于语音对话,实现端到端的自然交互。

Spoken Conversational Agents with Large Language Models

  • 从传统分步流程转向端到端语音-文本联合训练
  • 融合检索与视觉信息,提升对话上下文理解能力
  • 适合研究者与工程师快速上手语音助手开发

语音对话系统正朝着原生语音的大语言模型演进。本教程梳理了从分步式语音识别/自然语言理解到端到端、基于检索与视觉引导系统的演进路径。我们探讨了文本大模型向音频适配、跨模态对齐及语音-文本联合训练的方法;回顾了常用数据集、评估指标及在不同口音下的鲁棒性表现,并对比了系统设计选择(分步式与端到端、后处理纠错、流式处理)。本教程将工业级语音助手与当前开放域及任务导向型对话系统相连接,提出可复现的基线方法,并指出隐私、安全与评估方面的未解问题。参会者将获得实用开发方案与系统级发展路线图。

原文摘要 · Abstract (English)

Spoken conversational agents are converging toward voice-native LLMs. This tutorial distills the path from cascaded ASR/NLU to end-to-end, retrieval-and vision-grounded systems. We frame adaptation of text LLMs to audio, cross-modal alignment, and joint speech-text training; review datasets, metrics, and robustness across accents and compare design choices (cascaded vs. E2E, post-ASR correction, streaming). We link industrial assistants to current open-domain and task-oriented agents, highlight reproducible baselines, and outline open problems in privacy, safety, and evaluation. Attendees leave with practical recipes and a clear systems-level roadmap.

语音对话大模型端到端多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。