arXiv:2512.16832cs.CL2025-12ACL被引 2

语音语调比文字多传递十倍以上的情感和讽刺信息。

What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels

  • 用信息论方法量化语音与文字各自传递的意义信息量
  • 语音在情感和讽刺上比文字多传递超10倍信息
  • 适合研究语音理解、人机交互与跨模态分析的学者

语音语调(prosody)传达了大量文字无法体现的关键信息。本文提出一种基于信息论的方法,量化仅由语音传递而文字未表达的信息量,并探究这些信息的内容。通过大型语音与语言模型,计算特定语义维度(如情绪)与通信通道(如音频或文本)之间的互信息。我们利用电视和播客中的语音数据,测量了语音与文字在讽刺、情绪和疑问性方面的信息传递能力。结果表明,在缺乏长期上下文的情况下,语音通道在表达讽刺与情绪方面提供的信息量远超文字,超过一个数量级;而对于疑问性,语音提供的额外信息较少。文章最后提出将该方法扩展至更多语义维度、通信渠道及语言的研究计划。

原文摘要 · Abstract (English)

Prosody -- the melody of speech -- conveys critical information often not captured by the words or text of a message. In this paper, we propose an information-theoretic approach to quantify how much information is expressed by prosody alone and not by text, and crucially, what that information is about. Our approach applies large speech and language models to estimate the mutual information between a particular dimension of an utterance's meaning (e.g., its emotion) and any of its communication channels (e.g., audio or text). We then use this approach to quantify how much information is conveyed by audio and text about sarcasm, emotion, and questionhood, using speech from television and podcasts. We find that for sarcasm and emotion the audio channel -- and by implication the prosodic channel -- transmits over an order of magnitude more information about these features than the text channel alone, at least when long-term context beyond the current sentence is unavailable. For questionhood, prosody provides comparatively less additional information. We conclude by outlining a program applying our approach to more dimensions of meaning, communication channels, and languages.

语音分析信息论多模态情感识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。