用文本数据辅助训练,让语音对话系统跨领域更通用
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
- 联合训练语音与文本数据,提升跨领域泛化能力
- 无需目标领域语音数据,仍能实现优异跨域性能
- 适合资源有限、难以收集语音标注数据的场景
端到端语音对话状态追踪(DST)面临两大挑战:处理语音输入和数据稀缺。现有方法结合语音基础编码器与大语言模型虽取得良好效果,但在跨领域泛化上表现不佳,且需为每个目标领域提供标注的语音DST数据,而此类数据收集成本高、难度大。考虑到文本DST数据在多领域中更易获取,本文提出在已有语音DST数据和来自其他领域的文本DST数据上进行联合训练,以实现跨领域泛化。实验表明,该方法可在不依赖目标领域语音训练数据的情况下,获得良好的跨域性能。
原文摘要 · Abstract (English)
End-to-end spoken dialogue state tracking (DST) is made difficult by the tandem of having to handle speech input and data scarcity. Combining speech foundation encoders and large language models has been proposed in recent work as to alleviate some of this difficulty. Although this approach has been shown to result in strong spoken DST models, achieving state-of-the-art performance in realistic multi-turn DST, it struggles to generalize across domains and requires annotated spoken DST training data for each domain of interest. However, collecting such data for every target domain is both costly and difficult. Noting that textual DST data is more easily obtained for various domains, in this work, we propose jointly training on available spoken DST data and written textual data from other domains as a way to achieve cross-domain generalization. We conduct experiments which show the efficacy of our proposed method for getting good cross-domain DST performance without relying on spoken training data from the target domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。