用少量语音数据提升对话理解,让助手更懂人话。
Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling
- 设计辅助任务,用少量语音数据增强多模态理解。
- 仅用10%数据即达当前最优,性能超越主流模型。
- 适合语音对话系统研发者,尤其关注数据效率的团队。
对话助手在各类实际应用中日益普及,对先进多模态语音建模的需求也愈发迫切。语音作为自然交流方式,蕴含丰富的用户特征(如语速、音高),对有效交互至关重要。本文提出一种以数据为中心的定制化方法,高效提升对话语音建模中的多模态理解能力。核心是设计一种新型多任务学习范式,通过引入辅助任务,利用少量语音数据实现优化。该方法在Spoken-SQuAD基准上仅使用10%训练数据,便达到当前最优性能,构建了一个鲁棒高效的音频主导对话建模框架。此外,我们还推出了ASK-QA——首个包含多轮模糊用户请求与动态评估输入的口语对话数据集。代码与数据即将公开。
原文摘要 · Abstract (English)
Conversational assistants are increasingly popular across diverse real-world applications, highlighting the need for advanced multimodal speech modeling. Speech, as a natural mode of communication, encodes rich user-specific characteristics such as speaking rate and pitch, making it critical for effective interaction. Our work introduces a data-centric customization approach for efficiently enhancing multimodal understanding in conversational speech modeling. Central to our contributions is a novel multi-task learning paradigm that involves designing auxiliary tasks to utilize a small amount of speech data. Our approach achieves state-of-the-art performance on the Spoken-SQuAD benchmark, using only 10% of the training data with open-weight models, establishing a robust and efficient framework for audio-centric conversational modeling. We also introduce ASK-QA, the first dataset for multi-turn spoken dialogue with ambiguous user requests and dynamic evaluation inputs. Code and data forthcoming.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。