首个融合语音与情感标注的对话摘要数据集,助力更自然的语音摘要研究。
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
- 用大模型生成带填充词和回应语的对话脚本,再标注情绪、音高、语速。
- 构建13,460段语音对话,每段配事实摘要和情感摘要。
- 端到端语音模型比分步式系统情感摘要效果提升28%,适合语音理解研究者。
现有音频语言模型可处理长对话,但情感感知或口语对话摘要研究受限于缺乏同时包含语音、摘要和副语言线索的数据。本文提出Spoken DialogSum,首个将原始对话音频与事实摘要、情感丰富摘要及说话人年龄、性别、情绪等逐句标签对齐的语料库。数据构建分两阶段:首先,利用大模型重写DialogSum脚本,加入类似Switchboard的填充词与回应语,并为每句标注情绪、音高和语速;其次,使用富有表现力的语音合成引擎生成与副语言标签对齐的语音。该数据集包含13,460段情绪多样的对话,每段配有一个事实摘要和一个情感摘要。我们提供在线演示(https://fatfat-emosum.github.io/EmoDialog-Sum-Audio-Samples/),并计划近期发布完整数据集。基线实验表明,音频大模型在情感摘要上的ROUGE-L得分相比分步式ASR+LLM系统提升28%,验证了端到端语音建模的价值。
原文摘要 · Abstract (English)
Recent audio language models can follow long conversations. However, research on emotion-aware or spoken dialogue summarization is constrained by the lack of data that links speech, summaries, and paralinguistic cues. We introduce Spoken DialogSum, the first corpus aligning raw conversational audio with factual summaries, emotion-rich summaries, and utterance-level labels for speaker age, gender, and emotion. The dataset is built in two stages: first, an LLM rewrites DialogSum scripts with Switchboard-style fillers and back-channels, then tags each utterance with emotion, pitch, and speaking rate. Second, an expressive TTS engine synthesizes speech from the tagged scripts, aligned with paralinguistic labels. Spoken DialogSum comprises 13,460 emotion-diverse dialogues, each paired with both a factual and an emotion-focused summary. We release an online demo at https://fatfat-emosum.github.io/EmoDialog-Sum-Audio-Samples/, with plans to release the full dataset in the near future. Baselines show that an Audio-LLM raises emotional-summary ROUGE-L by 28% relative to a cascaded ASR-LLM system, confirming the value of end-to-end speech modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。