arXiv:2412.07116cs.LGcs.AI2024-12综述被引 23

综述生成技术在情绪合成中的应用,涵盖多模态与评估方法

A Review of Human Emotion Synthesis Based on Generative Technology

  • 系统梳理生成模型在表情、语音、文本等多模态情绪合成中的应用
  • 总结常用数据集与主流评价指标,揭示当前研究现状
  • 适合对情感计算和生成模型交叉领域感兴趣的科研人员

情感合成是情感计算的关键环节,旨在通过计算方法在多种模态中模拟和传达人类情绪,以实现更自然有效的人机交互。近年来,自编码器、生成对抗网络、扩散模型、大语言模型和序列到序列模型等生成模型的进展显著推动了该领域的发展。然而,该领域仍缺乏全面系统的综述。为此,本文旨在填补这一空白,提供基于生成模型的人类情绪合成的全面、系统性综述。首先介绍综述方法、涉及的情绪模型、生成模型的数学原理及使用数据集;随后覆盖不同生成模型在面部图像、语音和文本等多模态情绪合成中的应用;同时分析主流评估指标。此外,归纳主要发现并提出未来研究方向,为理解生成技术在复杂情绪合成中的作用提供全面视角。

原文摘要 · Abstract (English)

Human emotion synthesis is a crucial aspect of affective computing. It involves using computational methods to mimic and convey human emotions through various modalities, with the goal of enabling more natural and effective human-computer interactions. Recent advancements in generative models, such as Autoencoders, Generative Adversarial Networks, Diffusion Models, Large Language Models, and Sequence-to-Sequence Models, have significantly contributed to the development of this field. However, there is a notable lack of comprehensive reviews in this field. To address this problem, this paper aims to address this gap by providing a thorough and systematic overview of recent advancements in human emotion synthesis based on generative models. Specifically, this review will first present the review methodology, the emotion models involved, the mathematical principles of generative models, and the datasets used. Then, the review covers the application of different generative models to emotion synthesis based on a variety of modalities, including facial images, speech, and text. It also examines mainstream evaluation metrics. Additionally, the review presents some major findings and suggests future research directions, providing a comprehensive understanding of the role of generative technology in the nuanced domain of emotion synthesis.

情感合成生成模型多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。