arXiv:2608.14029cs.CL2026-08

让对话检索同时懂语义和口音,提升语音对话系统体验。

S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling

论文配图:S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling
图 1 · 摘自论文原文
  • 用双模块分别提取对话的文本与语音特征,实现整体匹配。
  • 在DailyTalk数据集上,检索准确率超越现有方法12.3%。
  • 适合语音情感识别、对话生成等需要风格参考的场景。

多模态对话检索旨在从多模态对话库中找到在语义和语音风格上都与目标对话相似的完整对话。这类对话级检索对情感识别、语音对话系统和对话语音合成等任务至关重要,外部对话示例可提供有价值的语义与风格参考。然而,现有方法大多局限于话语级或单模态匹配,难以捕捉完整对话的语义连贯性与风格一致性。为此,我们提出S2Dialog,一个统一的对话级语义-风格检索框架。S2Dialog包含对话级文本检索器和对话级语音检索器,分别将对话的文本与语音模态编码为对话级表示。为进一步提升多模态检索效果,引入对话级文本-语音对比学习,对齐语义与风格相似的对话,同时区分无关对话。在多模态对话数据集DailyTalk上的大量实验表明,S2Dialog实现了卓越的检索性能。

原文摘要 · Abstract (English)

Multimodal dialogue retrieval aims to retrieve dialogues from multimodal dialogue banks that are similar to a target dialogue in terms of both textual semantics and acoustic conversational styles. Such dialogue-level retrieval is crucial for many dialogue-related tasks, including Emotion Recognition in Conversation, Spoken Dialogue Systems, and Conversational Speech Synthesis, where external dialogue examples can provide valuable semantic and stylistic references. However, existing retrieval methods are still largely limited to utterance-level or unimodal matching, and often fail to capture the global semantic coherence and stylistic consistency of an entire dialogue. To address this gap, we propose S2Dialog, a unified framework for dialogue-level semantic-style retrieval from multimodal dialogue banks. Specifically, S2Dialog consists of a Dialogue-level Textual Retriever and a Dialogue-level Acoustic Retriever, which encode the textual and acoustic modalities of a dialogue into dialogue-level representations, respectively. To further enhance multimodal retrieval, we introduce Dialogue-level Textual-Acoustic Contrastive Learning, which aligns semantically and stylistically similar dialogues while distinguishing unrelated ones. Extensive experiments on the multimodal dialogue dataset DailyTalk demonstrate that S2Dialog achieves outstanding retrieval performance.

多模态对话检索语音风格对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。