通过语义与语调融合,提升语音中讽刺意图的识别效果。
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
- 用LLaMA 3模型捕捉话语级讽刺信号,结合真实讽刺语音样本提取语调特征。
- 联合使用语义与语调线索后,下游任务F1值最优,主观评分也最高。
- 适合研究情感语音、人机交互与讽刺识别的学者和开发者。
讽刺是一种语用现象,说话人传达的意义与字面内容相悖,依赖语义与语调表达的相互作用。然而,这些线索如何共同促成讽刺识别仍不明确。本文提出一种计算框架,将讽刺建模为语义理解与语调实现的融合。语义线索来自微调后的LLaMA 3模型,用于捕捉话语层面的讽刺意图;语调线索则从讽刺语音数据库中提取,基于语义对齐的语句获得讽刺语调实例。在语音合成测试平台上,感知评估显示语义与语调线索均能增强讽刺感知度,联合系统在下游任务中取得最佳F1值,同时保持高主观讽刺评分。结果表明,语义与语调在语用解读中具有互补作用,并展示了建模方法对揭示讽刺交流机制的启发意义。
原文摘要 · Abstract (English)
Sarcasm is a pragmatic phenomenon in which speakers convey meanings that diverge from literal content, relying on an interaction between semantics and prosodic expression. However, how these cues jointly contribute to the recognition of sarcasm remains poorly understood. We propose a computational framework that models sarcasm as the integration of semantic interpretation and prosodic realization. Semantic cues are derived from an LLaMA 3 model fine-tuned to capture discourse-level markers of sarcastic intent, while prosodic cues are extracted through semantically aligned utterances drawn from a database of sarcastic speech, providing prosodic exemplars of sarcastic delivery. Using a speech synthesis testbed, perceptual evaluations show that semantic and prosodic cues enhance perceived sarcasm, with the combined system achieving the best downstream F1 while maintaining high subjective sarcasm ratings. These findings highlight the complementary roles of semantics and prosody in pragmatic interpretation and illustrate how modeling can shed light on the mechanisms underlying sarcastic communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。