用可控语音技术研究讽刺语气的声调线索,发现音量影响最大。
What Makes Synthetic Speech Sound Sarcastic? A Prosody-Controlled Perception Study
- 用提示词控制语音速率、音高和音量,实现声调精细调节
- 人听者判断显示音量是讽刺感知主因,模型则更依赖语速
- 适合研究语音情感、人机交互或语音合成的学者参考
声调在讽刺感知中起关键作用,但以往研究依赖自然语音,难以分离各声学特征的独立影响。本文提出基于提示词的神经文本转语音框架,可分别调控语音速率、音高变化和响度,构建正交刺激集以实现因果测试。人类听众对讽刺程度和自然度进行评分,并与能处理音频的基础模型预测结果对比。结果显示,响度是人类感知讽刺的主要依据,而模型更重视语速,表明两者行为存在显著差异。该研究展示了可控神经TTS在语音感知机制研究中的潜力。
原文摘要 · Abstract (English)
Prosody plays an important role in sarcasm perception, yet previous studies have relied on naturally produced speech that lacks fine-grained control over individual acoustic dimensions. As prosodic cues co-vary in natural data, isolating their independent contributions remains challenging. We introduce a controlled framework using neural text-to-speech (TTS) with prompt-based prosodic conditioning to manipulate speech rate, pitch variation, and loudness. An orthogonal stimulus set was constructed to enable causal testing of prosodic cue effects. Human listeners rated sarcasm and naturalness, and their judgments were compared with predictions from a foundation model capable of processing audio input. Results show that loudness primarily drives human sarcasm perception, whereas the model assigns greater weight to speech rate, indicating limited behavioral alignment. This study shows how controllable neural TTS enables investigation of prosodic cue weighting in speech perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。