综述情绪识别与生成在表情、语音、文本中的最新进展
Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities
- 系统梳理三模态情绪分析的技术路径与理论基础
- 对比评估主流方法,揭示现有研究的局限性
- 适合初入该领域的研究者快速掌握全貌
情绪识别与生成已成为人工智能研究的重要方向,在医疗健康、客户服务等领域对人机交互具有重要意义。尽管已有部分关于情绪识别或生成的综述,但多数研究分散且局限于特定方法,缺乏对多模态近期发展与趋势的全面回顾。本文面向初学者,提供涵盖面部、语音和文本三模态的情绪识别与生成的综合性综述。介绍各模态背后的基本原理,将近年前沿研究按技术路径分类,并阐明其理论依据与动机,帮助理解实际应用。同时讨论评估指标、方法比较及当前挑战,提出未来研究方向,以推动更稳健、高效且符合伦理的情绪感知与生成系统的发展。
原文摘要 · Abstract (English)
Emotion recognition and generation have emerged as crucial topics in Artificial Intelligence research, playing a significant role in enhancing human-computer interaction within healthcare, customer service, and other fields. Although several reviews have been conducted on emotion recognition and generation as separate entities, many of these works are either fragmented or limited to specific methodologies, lacking a comprehensive overview of recent developments and trends across different modalities. In this survey, we provide a holistic review aimed at researchers beginning their exploration in emotion recognition and generation. We introduce the fundamental principles underlying emotion recognition and generation across facial, vocal, and textual modalities. This work categorises recent state-of-the-art research into distinct technical approaches and explains the theoretical foundations and motivations behind these methodologies, offering a clearer understanding of their application. Moreover, we discuss evaluation metrics, comparative analyses, and current limitations, shedding light on the challenges faced by researchers in the field. Finally, we propose future research directions to address these challenges and encourage further exploration into developing robust, effective, and ethically responsible emotion recognition and generation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。