arXiv:2606.06168cs.AIcs.CL2026-06中稿 · Interspeech 2026, …

通过语音语调不一致检测讽刺,比现有方法更准。

ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity

论文配图:ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity
图 1 · 摘自论文原文
  • 用全局情绪和时序语调双路径建模语调与情感基调的不匹配。
  • 在MUStARD++数据集上达到75.3的F1分数,跨语言也表现良好。
  • 能定位讽刺开始时间,且模型不确定度符合人类感知模糊性。

我们提出ProSarc,一种仅依赖音频的讽刺识别框架,通过建模时序语调不一致(即局部语调动态与整体情感基线之间的不匹配)来检测讽刺。该框架采用双编码路径:全局情绪编码器与时序语调编码器(BiLSTM + 多头注意力),共同输入到语调不一致分析模块,输出标量不一致得分用于分类。蒙特卡洛丢弃提供不确定性估计,基于注意力机制可定位讽刺起始点,无需帧级标注。ProSarc在MUStARD++上取得F1=75.3,泛化至自发口语(PodSarc,F1=62.9)和跨语言语音(MuSaG,F1=65.6)。十次运行验证表明不一致建模显著提升性能(Wilcoxon p=0.002,Cohen's d=1.51)。人工评估显示模型不确定性与感知模糊性一致,预测起始点与人工标注的时间窗口吻合。

原文摘要 · Abstract (English)

We present ProSarc, an audio-only framework that detects sarcasm by modelling temporal prosodic incongruity, that is, the mismatch between local prosodic dynamics and the utterance-level emotional baseline. Dual encoding paths, a Global Emotion Encoder and a Temporal Prosody Encoder (BiLSTM + multi-head attention), feed a Prosodic Incongruity Analyzer that produces a scalar incongruity score for classification. Monte Carlo dropout provides uncertainty estimates, and an attention-based mechanism localises sarcastic onset without frame-level labels. ProSarc outperforms prior audio-only methods on MUStARD++ (F1=75.3) and generalises to spontaneous (PodSarc, F1=62.9) and cross-lingual speech (MuSaG, F1=65.6). Ten-run validation confirms the contribution of incongruity modelling (Wilcoxon p=0.002, Cohen's d=1.51). Human evaluation shows that model uncertainty tracks perceptual ambiguity and predicted onsets align with human-annotated temporal windows.

语音分析讽刺识别语调建模多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。