提出语音的隐含信息理解框架,推动人机对话更自然
BoSS: Beyond-Semantic Speech
- 构建从指令识别到类人社交的语音系统能力层级
- 现有语音模型难以解析情感、上下文等超越语义的信息
- 适合研究人机交互与语音智能的学者和工程师
人类交流不仅包含显性语义,还依赖隐含信号与情境线索来塑造意义。然而,当前语音技术如自动语音识别(ASR)和文本转语音(TTS)常忽略这些超越语义的维度。为更好地表征与评估语音智能进展,我们提出语音交互系统能力等级(L1-L5),描绘了语音对话系统从基础命令识别到类人社交互动的演进路径。为此,我们引入超越语义语音(BoSS),指语音交流中超越显性语义的多维信息,包括情感、上下文动态与隐含语义,通过声调、节奏等特征扩展或修饰意义,增强对交际意图与场景的理解。我们建立基于认知相关性理论与机器学习模型的正式分析框架,评估五个维度的BoSS属性,发现当前语音语言模型(SLMs)在解析超越语义信号方面存在显著局限。结果凸显了推进BoSS研究的重要性,以实现更丰富、更情境感知的人机通信。
原文摘要 · Abstract (English)
Human communication involves more than explicit semantics, with implicit signals and contextual cues playing a critical role in shaping meaning. However, modern speech technologies, such as Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) often fail to capture these beyond-semantic dimensions. To better characterize and benchmark the progression of speech intelligence, we introduce Spoken Interaction System Capability Levels (L1-L5), a hierarchical framework illustrated the evolution of spoken dialogue systems from basic command recognition to human-like social interaction. To support these advanced capabilities, we propose Beyond-Semantic Speech (BoSS), which refers to the set of information in speech communication that encompasses but transcends explicit semantics. It conveys emotions, contexts, and modifies or extends meanings through multidimensional features such as affective cues, contextual dynamics, and implicit semantics, thereby enhancing the understanding of communicative intentions and scenarios. We present a formalized framework for BoSS, leveraging cognitive relevance theories and machine learning models to analyze temporal and contextual speech dynamics. We evaluate BoSS-related attributes across five different dimensions, reveals that current spoken language models (SLMs) are hard to fully interpret beyond-semantic signals. These findings highlight the need for advancing BoSS research to enable richer, more context-aware human-machine communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。