arXiv:2608.17931cs.CLcs.MM2026-08中稿 · ACM Multimedia 202…

构建细粒度语音情感数据集,提升对语气态度的识别能力

SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis

论文配图:SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis
图 1 · 摘自论文原文
  • 定义8类基于韵律的细微人际态度分类体系
  • 实验证明有声学信息的模型性能显著优于纯文本模型
  • 适合需要理解说话人语气的社会交互场景研究

近年来人工智能在语音处理领域取得突破,但有效理解语音不仅需知其言,更需察其声。语音情感分析在招聘、客户服务等场景中至关重要,但现有研究存在两大局限:一是主流方法依赖文本中心流程,经语音识别后分析文本,丢失了韵律、语调等关键声学特征,难以捕捉声学模糊语句中的态度信息;二是现有评测基准标签粒度粗,侧重基础情绪(如高兴、悲伤),忽略自信、不耐烦等社交敏感的细微人际立场。为此,我们提出SpeechSense数据集,聚焦细粒度语音情感分析。具体地,定义了以韵律为主要依据的8类人际立场分类体系,并基于高保真语音合成与严格人工验证构建数据集。多模态大模型、纯文本大模型及语音编码器的对比实验表明,具备声学输入的模型性能持续优于仅用文本的基线。结果实证了声学线索在识别微妙说话者态度中的主导作用,凸显SpeechSense的价值。数据集及补充材料见https://github.com/Sher13cked/SpeechSense。

原文摘要 · Abstract (English)

Recent advances in AI have revolutionized speech processing, yet effective speech understanding requires discerning not just what is said, but how it is said. Speech Sentiment Analysis plays a critical role in decoding these paralinguistic cues for diverse real-world applications such as recruitment and customer service. However, existing Speech Sentiment Analysis research faces two primary limitations. First, dominant approaches rely on text-centric pipelines that cascade Automatic Speech Recognition with text analysis. This process inevitably discards essential acoustic features like prosody and tone, failing to capture attitudinal meanings in acoustically ambiguous utterances. Second, current benchmarks suffer from a mismatch in label granularity, prioritizing basic emotions (e.g., happy, sad) over the nuanced interpersonal stances (e.g., confident, impatient) necessary for social sensitivity. To address these limitations, we propose a novel dataset, SpeechSense, for fine-grained speech sentiment analysis. Specifically, we define a specialized 8-class taxonomy of interpersonal stances detectable primarily through prosodic cues beyond lexical content alone. We then construct a curated dataset based on this taxonomy, built from high-fidelity speech synthesis and rigorous human validation. Comprehensive experiments across multi-modal LLMs, text-only LLMs, and speech encoders demonstrate that models with acoustic access consistently outperform text-only baselines. These results empirically validate the primacy of acoustic cues in detecting subtle speaker attitudes, highlighting the necessity of SpeechSense. Dataset and supplementary materials are available at https://github.com/Sher13cked/SpeechSense.

语音情感分析细粒度标注声学特征韵律分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。