用合成语音提升认知状态识别,效果优于纯文本模型。
Synthetic Audio Helps for Cognitive State Tasks
- 用现成语音合成模型生成零样本音频,与文本联合训练。
- 7项认知状态任务均因加入合成音频提升性能,部分接近真实音频效果。
- 适合关注多模态情感分析、语音辅助认知建模的研究者。
自然语言处理领域长期聚焦于纯文本的认知状态任务,但语音中的语调信息能提供关键补充。我们提出,语音合成(TTS)模型在生成自然语音过程中会隐式捕捉认知状态特征,而这些特征与语言模型所利用的信息正交。为此,我们提出合成音频微调框架(SAD),在7个认知状态相关任务中证明:将文本与来自现成TTS系统的零样本合成音频联合训练,可显著提升性能。在无真实音频的语料上,加入合成音频优于纯文本基线;在包含真实音频的任务中,使用合成音频的SAD框架表现可媲美使用真实音频的模型。
原文摘要 · Abstract (English)
The NLP community has broadly focused on text-only approaches of cognitive state tasks, but audio can provide vital missing cues through prosody. We posit that text-to-speech models learn to track aspects of cognitive state in order to produce naturalistic audio, and that the signal audio models implicitly identify is orthogonal to the information that language models exploit. We present Synthetic Audio Data fine-tuning (SAD), a framework where we show that 7 tasks related to cognitive state modeling benefit from multimodal training on both text and zero-shot synthetic audio data from an off-the-shelf TTS system. We show an improvement over the text-only modality when adding synthetic audio data to text-only corpora. Furthermore, on tasks and corpora that do contain gold audio, we show our SAD framework achieves competitive performance with text and synthetic audio compared to text and gold audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。