用语音大模型实现零样本槽位填充,统一处理语音与语义理解。
SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
- 基于语音大模型构建端到端的槽位填充框架,支持零样本泛化。
- 在多个数据集上显著提升性能,尤其在未见槽位标签上表现优异。
- 为实际应用提供训练数据、架构和策略优化的实证指导。
槽位填充是语音理解中的关键任务,传统方法依赖语音识别与自然语言理解的级联结构。近期出现的语音大语言模型(SpeechLLMs)融合语音与文本基础模型,为实现更统一、生成式、指令跟随的语音理解提供了新路径,同时具备零样本能力,能泛化至未见过的槽位标签。本文通过构建该任务的实验上界,识别出性能、鲁棒性与泛化能力的差距,并提出改进训练数据、模型架构与训练策略的方法以缩小差距。结果表明,各项改进均显著提升性能,同时揭示了实际应用中的挑战,为有效利用这类新兴模型提供了实证依据与实践指导。
原文摘要 · Abstract (English)
Slot filling is a crucial subtask in spoken language understanding (SLU), traditionally implemented as a cascade of speech recognition followed by one or more natural language understanding (NLU) components. The recent advent of speech-based large language models (speechLLMs), which integrate speech and textual foundation models, has opened new avenues for achieving speech understanding tasks in a more unified, generative, and instruction-following manner while promising data and compute efficiency with zero-shot abilities, generalizing to unseen slot labels. We address the slot-filling task by creating an empirical upper bound for the task, identifying performance, robustness, and generalization gaps, and proposing improvements to the training data, architecture, and training strategies to narrow the gap with the upper bound result. We show that each of these measures improve performance substantially, while highlighting practical challenges and providing empirical guidance and insights for harnessing these emerging models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。