用合成数据提升对话系统应对复杂场景的能力
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
- 构建大规模跨场景语音对话数据集ShareChatX,覆盖音乐、情感等多样情境
- 提出OmniChat模型,通过异构特征融合实现多轮对话中关键信息的动态选择
- 验证合成数据与真实数据的最佳配比,显著提升真实场景表现
随着大语言模型的发展,语音对话系统已能实现自然人机交互,但仍难以应对真实对话中的复杂性,如音频事件、音乐背景和情感表达。主要原因在于现有对话数据集在规模和场景多样性上受限。本文提出利用合成数据增强对话模型在多样化场景下的能力。我们构建了首个全面且大规模的语音对话数据集ShareChatX,涵盖多种现实场景。基于此,提出OmniChat系统,其包含异构特征融合模块,可针对不同对话上下文优化特征选择。通过大量实验,确定了合成数据与真实数据的理想比例,在真实对话数据集DailyTalk上达到当前最优性能。研究强调了合成数据在处理涉及音频与音乐等复杂场景中的关键作用。
原文摘要 · Abstract (English)
With the rapid development of large language models, researchers have created increasingly advanced spoken dialogue systems that can naturally converse with humans. However, these systems still struggle to handle the full complexity of real-world conversations, including audio events, musical contexts, and emotional expressions, mainly because current dialogue datasets are constrained in both scale and scenario diversity. In this paper, we propose leveraging synthetic data to enhance the dialogue models across diverse scenarios. We introduce ShareChatX, the first comprehensive, large-scale dataset for spoken dialogue that spans diverse scenarios. Based on this dataset, we introduce OmniChat, a multi-turn dialogue system with a heterogeneous feature fusion module, designed to optimize feature selection in different dialogue contexts. In addition, we explored critical aspects of training dialogue systems using synthetic data. Through comprehensive experimentation, we determined the ideal balance between synthetic and real data, achieving state-of-the-art results on the real-world dialogue dataset DailyTalk. We also highlight the crucial importance of synthetic data in tackling diverse, complex dialogue scenarios, especially those involving audio and music. For more details, please visit our demo page at \url{https://sharechatx.github.io/}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。