通过语音语调不一致检测讽刺,比现有方法更准。
ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity

- 用全局情绪和时序语调双路径建模语调与情感基调的不匹配。
- 在MUStARD++数据集上达到75.3的F1分数,跨语言也表现良好。
- 能定位讽刺开始时间,且模型不确定度符合人类感知模糊性。
我们提出ProSarc,一种仅依赖音频的讽刺识别框架,通过建模时序语调不一致(即局部语调动态与整体情感基线之间的不匹配)来检测讽刺。该框架采用双编码路径:全局情绪编码器与时序语调编码器(BiLSTM + 多头注意力),共同输入到语调不一致分析模块,输出标量不一致得分用于分类。蒙特卡洛丢弃提供不确定性估计,基于注意力机制可定位讽刺起始点,无需帧级标注。ProSarc在MUStARD++上取得F1=75.3,泛化至自发口语(PodSarc,F1=62.9)和跨语言语音(MuSaG,F1=65.6)。十次运行验证表明不一致建模显著提升性能(Wilcoxon p=0.002,Cohen's d=1.51)。人工评估显示模型不确定性与感知模糊性一致,预测起始点与人工标注的时间窗口吻合。
原文摘要 · Abstract (English)
We present ProSarc, an audio-only framework that detects sarcasm by modelling temporal prosodic incongruity, that is, the mismatch between local prosodic dynamics and the utterance-level emotional baseline. Dual encoding paths, a Global Emotion Encoder and a Temporal Prosody Encoder (BiLSTM + multi-head attention), feed a Prosodic Incongruity Analyzer that produces a scalar incongruity score for classification. Monte Carlo dropout provides uncertainty estimates, and an attention-based mechanism localises sarcastic onset without frame-level labels. ProSarc outperforms prior audio-only methods on MUStARD++ (F1=75.3) and generalises to spontaneous (PodSarc, F1=62.9) and cross-lingual speech (MuSaG, F1=65.6). Ten-run validation confirms the contribution of incongruity modelling (Wilcoxon p=0.002, Cohen's d=1.51). Human evaluation shows that model uncertainty tracks perceptual ambiguity and predicted onsets align with human-annotated temporal windows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。