首篇系统综述语音讽刺识别,揭示多模态融合的潜力与挑战。
Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects
- 首次聚焦语音讽刺识别,从单模态向多模态融合演进
- 现有语音讽刺数据集稀缺,特征提取已从传统声学转向深度学习表示
- 适合关注人机交互、跨文化语言理解的研究者
讽刺是人际与人机交互中的常见现象,其识别对提升沟通质量至关重要。语言学研究指出,语调、语速和音高变化等语音线索在表达讽刺意图中起关键作用。尽管已有研究集中于文本讽刺检测,但语音数据在讽刺识别中的价值仍被低估。近年来,语音技术的进步凸显了利用语音数据实现自动讽刺识别的重要性,有助于改善神经退行性疾病患者的社交互动,并推动机器对复杂人类语言的深层理解。本系统综述首次聚焦语音驱动的讽刺识别,梳理了从单模态到多模态方法的演进历程,涵盖数据集、特征提取与分类方法,旨在弥合不同研究领域间的鸿沟。研究发现:语音讽刺数据集仍显不足;特征提取技术已由传统声学特征发展为深度学习表示;分类方法则从单模态走向多模态融合。由此提出需加强跨文化和多语言讽刺识别研究,并将讽刺视为多模态现象而非仅文本问题。
原文摘要 · Abstract (English)
Sarcasm, a common feature of human communication, poses challenges in interpersonal interactions and human-machine interactions. Linguistic research has highlighted the importance of prosodic cues, such as variations in pitch, speaking rate, and intonation, in conveying sarcastic intent. Although previous work has focused on text-based sarcasm detection, the role of speech data in recognizing sarcasm has been underexplored. Recent advancements in speech technology emphasize the growing importance of leveraging speech data for automatic sarcasm recognition, which can enhance social interactions for individuals with neurodegenerative conditions and improve machine understanding of complex human language use, leading to more nuanced interactions. This systematic review is the first to focus on speech-based sarcasm recognition, charting the evolution from unimodal to multimodal approaches. It covers datasets, feature extraction, and classification methods, and aims to bridge gaps across diverse research domains. The findings include limitations in datasets for sarcasm recognition in speech, the evolution of feature extraction techniques from traditional acoustic features to deep learning-based representations, and the progression of classification methods from unimodal approaches to multimodal fusion techniques. In so doing, we identify the need for greater emphasis on cross-cultural and multilingual sarcasm recognition, as well as the importance of addressing sarcasm as a multimodal phenomenon, rather than a text-based challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。