用声音和提示校准提升大模型讽刺检测准确率
CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection

- 通过双向提示校准消除大模型的讽刺偏见
- 结合声学特征使弱模型性能提升38.2%
- 适合跨语言讽刺识别与可解释性研究
讽刺检测是情感计算中的基础任务,但零样本指令微调的大语言模型在全能力范围内系统性高估讽刺类。本文提出无需微调的CHARM框架,包含双向电荷校准(BiCAL)模块,通过对称提示引导模型产生相反判断,构建的偏差相互抵消,实现无偏语用信号恢复;以及声学晚期融合救援(ALFR)模块,利用浅层分类器融合校准结果、声学特征与语音感知探针,主动降低饱和文本投票权重,增强声学证据影响。BiCAL在MUStARD上实现0.787的最高零样本文本仅宏F1,ALFR使弱模型在CMMA上性能提升高达+0.382。元分析显示在两数据集上均具极显著性(Z=13.89, Z=34.64;p<10^-43)。分析还发现低层声学特征跨语言迁移失败,高层感知抽象仍具鲁棒性。整体框架实现可解释、跨语言多模态讽刺检测。
原文摘要 · Abstract (English)
Sarcasm detection, the identification of discrepancies between literal and intended meaning, is a fundamental task in affective computing. However, zero-shot instruction-tuned Large Language Models (LLMs) systematically over-predict the positive (sarcastic) class across the entire capability spectrum, while the prosodic cues humans rely on remain underexploited and transfer unevenly across languages. We introduce CHARM (Charge Calibration and Acoustic Rescue for Multimodal Sarcasm Detection), a training-free framework that couples two modules. Bidirectional Charge Calibration (BiCAL) steers the LLM toward opposing sarcastic and literal verdicts along a symmetric axis of charged prompts; the induced directional biases cancel by construction, and a simple aggregation recovers an unbiased pragmatic signal. Acoustic Late-Fusion Rescue (ALFR) then fuses the calibrated votes with prosodic descriptors and LLM-generated auditory-perception probes through a shallow classifier, actively down-weighting saturated text votes in favour of acoustic evidence. Without fine-tuning any backbone, BiCAL attains the highest reported zero-shot text-only Macro-F1 of 0.787 on MUStARD, while ALFR lifts weak backbones by up to +0.382 Macro-F1 on CMMA. A Stouffer meta-analysis confirms statistical significance on MUStARD and CMMA (Z = 13.89 and Z = 34.64, respectively; p < 10^-43). Our analysis further uncovers a cross-cultural prosodic decoupling: low-level acoustics fail to transfer across languages, whereas high-level perceptual abstractions remain robust. Together, these components yield an explainable, cross-lingual multimodal detector.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。