arXiv:2410.12028cs.SDcs.LG2024-10被引 2

用情绪信息生成音频描述数据,提升语音字幕质量

EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation

  • 用声景情绪估计结果指导大模型生成带情绪标签的音频描述
  • 构建12万条含情绪信息的音频字幕数据集,提升描述准确性
  • 适合做音频理解、情感计算与多模态生成的研究者参考

近期音频-语言建模进展,如自动音频字幕生成,得益于大型语言模型生成的合成数据。然而,现有环境声音字幕方法主要关注音频事件标签,未充分利用录音中可能存在的情绪信息。本文提出通过向ChatGPT提供声景情绪识别(SER)估计信息,生成带情绪增强的合成音频字幕数据。我们构建了名为EmotionCaps的数据集,包含约12万条音频片段及其配对的合成描述,其中融合了声景情绪信息。我们假设这些额外的情绪信息能生成更贴合音频情感基调的高质量字幕,从而提升训练模型的表现。通过客观与主观评估对比多个基线模型,实验结果验证了该假设,挑战了当前字幕生成范式,并为未来模型开发与评估提供了新方向。

原文摘要 · Abstract (English)

Recent progress in audio-language modeling, such as automated audio captioning, has benefited from training on synthetic data generated with the aid of large-language models. However, such approaches for environmental sound captioning have primarily focused on audio event tags and have not explored leveraging emotional information that may be present in recordings. In this work, we explore the benefit of generating emotion-augmented synthetic audio caption data by instructing ChatGPT with additional acoustic information in the form of estimated soundscape emotion. To do so, we introduce EmotionCaps, an audio captioning dataset comprised of approximately 120,000 audio clips with paired synthetic descriptions enriched with soundscape emotion recognition (SER) information. We hypothesize that this additional information will result in higher-quality captions that match the emotional tone of the audio recording, which will, in turn, improve the performance of captioning models trained with this data. We test this hypothesis through both objective and subjective evaluation, comparing models trained with the EmotionCaps dataset to multiple baseline models. Our findings challenge current approaches to captioning and suggest new directions for developing and assessing captioning models.

音频字幕情绪识别合成数据多模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。