arXiv:2412.15995cs.CLcs.AI2024-12被引 3

用少量语音数据提升对话理解,让助手更懂人话。

Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling

  • 设计辅助任务,用少量语音数据增强多模态理解。
  • 仅用10%数据即达当前最优,性能超越主流模型。
  • 适合语音对话系统研发者,尤其关注数据效率的团队。

对话助手在各类实际应用中日益普及,对先进多模态语音建模的需求也愈发迫切。语音作为自然交流方式,蕴含丰富的用户特征(如语速、音高),对有效交互至关重要。本文提出一种以数据为中心的定制化方法,高效提升对话语音建模中的多模态理解能力。核心是设计一种新型多任务学习范式,通过引入辅助任务,利用少量语音数据实现优化。该方法在Spoken-SQuAD基准上仅使用10%训练数据,便达到当前最优性能,构建了一个鲁棒高效的音频主导对话建模框架。此外,我们还推出了ASK-QA——首个包含多轮模糊用户请求与动态评估输入的口语对话数据集。代码与数据即将公开。

原文摘要 · Abstract (English)

Conversational assistants are increasingly popular across diverse real-world applications, highlighting the need for advanced multimodal speech modeling. Speech, as a natural mode of communication, encodes rich user-specific characteristics such as speaking rate and pitch, making it critical for effective interaction. Our work introduces a data-centric customization approach for efficiently enhancing multimodal understanding in conversational speech modeling. Central to our contributions is a novel multi-task learning paradigm that involves designing auxiliary tasks to utilize a small amount of speech data. Our approach achieves state-of-the-art performance on the Spoken-SQuAD benchmark, using only 10% of the training data with open-weight models, establishing a robust and efficient framework for audio-centric conversational modeling. We also introduce ASK-QA, the first dataset for multi-turn spoken dialogue with ambiguous user requests and dynamic evaluation inputs. Code and data forthcoming.

语音对话多模态数据效率小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。