arXiv:2410.13385eess.AScs.CL2024-10被引 1

融合语音特征提升对话策略,尤其在语音转写差时表现更优

On the Use of Audio to Improve Dialogue Policies

  • 用双多头注意力融合语音与文本嵌入,增强对话决策能力
  • 在DSTC2数据集上用户需求得分提升9.8%相对性能
  • 适合语音识别不稳定的场景,对语音信息敏感的系统有参考价值

随着语音技术的快速发展,基于语音的目标导向对话系统日益普及。对话系统的核心模块之一是对话策略,通常仅依赖语音转录文本,严重受转录质量影响,并忽略语音中蕴含的重要非语言信息。本文提出新架构,通过双多头注意力机制融合语音与文本嵌入。实验表明,引入音频嵌入的对话策略优于纯文本方法,尤其在转写噪声较大的情况下;且文本与音频嵌入的融合方式对性能提升至关重要。在DSTC2数据集上,相比仅使用文本的系统,用户请求得分实现了9.8%的相对提升。

原文摘要 · Abstract (English)

With the significant progress of speech technologies, spoken goal-oriented dialogue systems are becoming increasingly popular. One of the main modules of a dialogue system is typically the dialogue policy, which is responsible for determining system actions. This component usually relies only on audio transcriptions, being strongly dependent on their quality and ignoring very important extralinguistic information embedded in the user's speech. In this paper, we propose new architectures to add audio information by combining speech and text embeddings using a Double Multi-Head Attention component. Our experiments show that audio embedding-aware dialogue policies outperform text-based ones, particularly in noisy transcription scenarios, and that how text and audio embeddings are combined is crucial to improve performance. We obtained a 9.8% relative improvement in the User Request Score compared to an only-text-based dialogue system on the DSTC2 dataset.

对话系统语音融合注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。