arXiv:2508.01960cs.SDcs.CL2025-08

研究非语言发声的表达与挑战,推动真实场景下的情感建模。

Non-Verbal Vocalisations and their Challenges: Emotion, Privacy, Sparseness, and Real Life

  • 提出基于语料库的建模方法,提升非语言发声的真实性。
  • 指出真实场景中数据稀疏且隐私受限,导致现有数据不足。
  • 适合情感计算、人机交互领域研究者参考。

非语言发声(NVVs)是无明确语义的短促非词发声,传达情绪或副语言信息。本文回顾了过去两个世纪心理学与语言学对这类发声的研究历程,指出其曾被忽视,后随情感研究兴起而重获关注。文章梳理了NVVs的类型特征与功能,以典型发声‘ah’为例说明。然而,该领域面临多重挑战:真实生活场景中的隐私问题限制了数据采集;当前多依赖孤立或刻意模仿的样本,难以反映真实语境。为此,本文主张采用语料库方法以实现更真实的建模,但仍需应对数据稀疏与隐私保护难题。

原文摘要 · Abstract (English)

Non-Verbal Vocalisations (NVVs) are short `non-word' utterances without proper linguistic (semantic) meaning but conveying connotations -- be this emotions/affects or other paralinguistic information. We start this contribution with a historic sketch: how they were addressed in psychology and linguistics in the last two centuries, how they were neglected later on, and how they came to the fore with the advent of emotion research. We then give an overview of types of NVVs (formal aspects) and functions of NVVs, exemplified with the typical NVV \textit{ah}. Interesting as they are, NVVs come, however, with a bunch of challenges that should be accounted for: Privacy and general ethical considerations prevent them of being recorded in real-life (private) scenarios to a sufficient extent. Isolated, prompted (acted) exemplars do not necessarily model NVVs in context; yet, this is the preferred strategy so far when modelling NVVs, especially in AI. To overcome these problems, we argue in favour of corpus-based approaches. This guarantees a more realistic modelling; however, we are still faced with privacy and sparse data problems.

非语言发声情感计算数据稀疏隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。