提升文本转音频模型对声音事件关系的理解能力
RiTTA: Modeling Event Relations in Text-to-Audio Generation
- 构建涵盖真实场景的音频事件关系语料库
- 提出新评估指标,多角度衡量关系建模效果
- 设计微调框架,增强现有模型的关系建模能力
尽管文本转音频(TTA)生成模型在高保真音频和细粒度上下文理解方面取得显著进展,但在建模输入文本中描述的声音事件间关系方面仍存在困难。以往的TTA方法未系统探索音频事件关系建模,也未提出相应增强框架。本文首次系统研究了TTA生成模型中的音频事件关系建模问题:1. 提出一个全面覆盖现实场景中所有潜在关系的关系语料库;2. 构建包含常见音频的音频事件语料库;3. 提出新的评估指标,从多个角度评估音频事件关系建模能力。此外,提出一种微调框架,以增强现有TTA模型对音频事件关系的建模能力。代码已开源。
原文摘要 · Abstract (English)
Despite significant advancements in Text-to-Audio (TTA) generation models achieving high-fidelity audio with fine-grained context understanding, they struggle to model the relations between audio events described in the input text. However, previous TTA methods have not systematically explored audio event relation modeling, nor have they proposed frameworks to enhance this capability. In this work, we systematically study audio event relation modeling in TTA generation models. We first establish a benchmark for this task by: 1. proposing a comprehensive relation corpus covering all potential relations in real-world scenarios; 2. introducing a new audio event corpus encompassing commonly heard audios; and 3. proposing new evaluation metrics to assess audio event relation modeling from various perspectives. Furthermore, we propose a finetuning framework to enhance existing TTA models ability to model audio events relation. Code is available at: https://github.com/yuhanghe01/RiTTA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。