arXiv:2604.01155cs.SD2026-04ACL被引 4

让音频语言模型同时理解片段和帧级内容,提升细节感知能力。

FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining

  • 用双流损失与聚类采样,融合粗粒度与细粒度标注数据。
  • 在多个任务上达到当前最优,尤其在声音事件检测中提升显著。
  • 适合需要精准时序定位的音频理解研究者使用。

对比预训练的音视频模型(如CLAP)在片段级理解上表现优异,但在帧级任务上表现不足。现有方法未能充分利用真实世界音视频数据中粗细粒度共存的特点——大量片段级文本描述与少量帧级标注并存。本文提出细粒度音视频预训练框架FineLAP,通过引入双流Sigmoid损失与基于聚类的采样策略,联合优化片段与帧级对齐。为捕捉全局语义与局部细节,采用解耦音频投影器配合自监督编码器。针对时间标注数据稀缺问题,构建了大规模合成声学事件检测数据集FineLAP-100k,通过可扩展的清洗流程生成。大量实验表明,FineLAP在检索、分类、声音事件检测及文到音定位等任务上均达当前最优性能。消融实验进一步证明粗粒度与细粒度对齐相互促进,为构建更优音视频模型提供新思路。

原文摘要 · Abstract (English)

Contrastively pretrained audio-language models (e.g., CLAP) excel at clip-level understanding but struggle with frame-level tasks. Existing extensions fail to exploit the varying granularity of real-world audio-text data, where massive clip-level textual descriptions coexist with limited frame-level annotations. This paper proposes Fine-grained Language-Audio Pretraining (FineLAP), a novel training paradigm that advances both clip- and frame-level alignment in CLAP with heterogeneous data. FineLAP introduces a dual-stream sigmoid loss with a cluster-based sampling strategy to jointly learn from clip- and frame-level supervision. To capture both global semantics and local details, FineLAP uses a decoupled audio projector on top of a self-supervised encoder. To alleviate the scarcity of temporally annotated data, we present FineLAP-100k, a large-scale synthetic SED dataset constructed through a scalable curation pipeline. Extensive experiments demonstrate that FineLAP achieves SOTA performance across multiple audio understanding tasks, including retrieval, classification, sound event detection, and text-to-audio grounding. Ablation studies further show that coarse- and fine-grained alignment are mutually beneficial, providing insights for building better audio-language models (ALMs).

音视频预训练细粒度对齐多粒度学习声音事件检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。