arXiv:2604.26417cs.CLcs.SD2026-04

让语音描述能捕捉情绪变化,提升人机交互的情感智能。

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses

论文配图:EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses
图 1 · 摘自论文原文
  • 构建动态情绪感知的语音描述框架,融合时序情绪变化与话语级语义。
  • 首个大规模话语级情绪转换数据集,支持情绪过渡识别与说话人分离。
  • 适配大模型生成描述与指令两种标注,助力情感化对话系统研发。

情绪感知与适应性表达是人机交互的核心能力。尽管语音情绪描述(SEC)已实现细粒度情绪建模,现有系统仍局限于孤立句子中的静态单情绪表征,忽视话语层面的情绪动态演变。为此,我们提出情绪转换感知的语音描述(EmoTransCap)范式,将时间维度情绪动态与话语级语音描述相结合。为构建富含情绪转换且可扩展的数据集,设计自动化数据构建流水线,形成首个专为捕捉话语级情绪转换而设计的大规模数据集。通过融合话语级语音的声学特征与时间线索,生成语义丰富的描述。提出的多任务情绪转换识别(MTETR)模型可联合完成情绪转换检测与说话人分离。借助大语言模型的语义分析能力,生成描述型与指令型两种标注版本。该数据集支持捕捉情绪转换的语音描述,促进时序动态、细粒度情绪理解。此外,还引入可控的、转换感知的话语级情绪语音合成系统,增强类人情感表达,推动情感智能对话代理发展。

原文摘要 · Abstract (English)

Emotion perception and adaptive expression are fundamental capabilities in human-agent interaction. While recent advances in speech emotion captioning (SEC) have improved fine-grained emotional modeling, existing systems remain limited to static, single-emotion characterization within isolated sentences, neglecting dynamic emotional transitions at the discourse level. To address this gap, we propose Emotion Transition-Aware Speech Captioning (EmoTransCap), a paradigm that integrates temporal emotion dynamics with discourse-level speech description. To construct a dataset rich in emotion transitions while enabling scalable expansion, we design an automated pipeline for dataset creation. This is the first large-scale dataset explicitly designed to capture discourse-level emotion transitions. To generate semantically rich descriptions, we incorporate acoustic attributes and temporal cues from discourse-level speech. Our Multi-Task Emotion Transition Recognition (MTETR) model performs joint emotion transition detection and diarization. Leveraging the semantic analysis capabilities of LLMs, we produce two annotation versions: descriptive and instruction-oriented. These data and annotations offer a valuable resource for advancing emotion perception and emotional expressiveness. The dataset enables speech captions that capture emotional transitions, facilitating temporal-dynamic and fine-grained emotion understanding. We also introduce a controllable, transition-aware emotional speech synthesis system at the discourse level, enhancing anthropomorphic emotional expressiveness and supporting emotionally intelligent conversational agents.

语音描述情绪识别对话系统大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。