arXiv:2505.07609eess.AScs.LG2025-05被引 20

用段落级音频描述提升音视频对齐精度,适合语音生成与检索任务。

TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining

  • 构建1.2万条带时间戳的音频-文本对,实现帧级语义对齐。
  • 在AudioSet强标签测试中,帧级对齐效果优于全局描述模型。
  • 使用大模型清洗标注,减少非听觉内容和语言偏见干扰。

将音频与文本描述关联对多种任务至关重要,包括预训练、零样本分类、音频检索、音频字幕生成及文本控制的音频生成。现有对比语言-音频预训练模型通常使用全局片段级描述进行训练,仅提供弱时间监督。我们假设,若希望模型输出帧级嵌入,则基于CLAP的模型可从更强的时间监督中获益。为验证该假设,我们从Freesound收集约1.2万条音频,每条均配有与特定时间片段对应的单句自由文本描述。利用大语言模型清洗标注,去除对不可听事件、转录语音、拼写错误及标注者语言偏见的提及。进一步提出一种帧级对比训练策略,学习将文本描述与音频中的时间区域对齐。在AudioSet Strong基准上的评估表明,相比仅使用全局描述训练的模型,本模型具有更优的时间文本-音频对齐能力。数据集和源代码分别公开于Zenodo和GitHub。

原文摘要 · Abstract (English)

Learning to associate audio with textual descriptions is valuable for a range of tasks, including pretraining, zero-shot classification, audio retrieval, audio captioning, and text-conditioned audio generation. Existing contrastive language-audio pretrained models are typically trained using global, clip-level descriptions, which provide only weak temporal supervision. We hypothesize that CLAP-like language-audio models - particularly, if they are expected to produce frame-level embeddings - can benefit from a stronger temporal supervision. To confirm our hypothesis, we curate a novel dataset of approximately 12,000 audio recordings from Freesound, each annotated with single-sentence free-text descriptions linked to a specific temporal segment in an audio recording. We use large language models to clean these annotations by removing references to non-audible events, transcribed speech, typos, and annotator language bias. We further propose a frame-wise contrastive training strategy that learns to align text descriptions with temporal regions in an audio recording and demonstrate that our model has better temporal text-audio alignment abilities compared to models trained only on global captions when evaluated on the AudioSet Strong benchmark. The dataset and our source code are available on Zenodo and GitHub, respectively.

音视频对齐语言模型音频标注对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。