研究对话中相同词语的语调相似性,发现自监督模型更贴近人耳感知。
Representation of perceived prosodic similarity of conversational feedback
- 通过三元比较实验测量语调相似性感知
- 自监督语音表征比基线特征更贴近人类感知
- 对比学习可进一步对齐表征与人类判断
话语反馈(如‘嗯’、‘是’、‘好’)是口语对话的重要组成部分,对确保对话系统中的共同认知至关重要。其确切含义由词汇形式和语调共同传达。本文研究具有相同词汇形式的话语反馈的语调感知相似性,并评估现有语音表征在反映这种相似性方面的表现。通过招募参与者进行三元比较任务,测量来自两个不同数据集的反馈响应的感知相似性。结果表明,频谱特征和自监督语音表征比提取的基频特征更能编码语调信息,尤其在同说话人反馈中表现更优。此外,通过对比学习可进一步压缩并对齐表征以逼近人类感知。
原文摘要 · Abstract (English)
Vocal feedback (e.g., `mhm', `yeah', `okay') is an important component of spoken dialogue and is crucial to ensuring common ground in conversational systems. The exact meaning of such feedback is conveyed through both lexical and prosodic form. In this work, we investigate the perceived prosodic similarity of vocal feedback with the same lexical form, and to what extent existing speech representations reflect such similarities. A triadic comparison task with recruited participants is used to measure perceived similarity of feedback responses taken from two different datasets. We find that spectral and self-supervised speech representations encode prosody better than extracted pitch features, especially in the case of feedback from the same speaker. We also find that it is possible to further condense and align the representations to human perception through contrastive learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。