arXiv:2505.05335cs.SDeess.AS2025-05ICML被引 26

让AI精准定位声音事件发生时间,支持未知声音类型

FLAM: Frame-Wise Language-Audio Modeling

  • 采用帧级对比学习与逻辑调整,解决训练中的虚假关联问题
  • 在大规模多事件数据集上实现高精度声音事件定位
  • 适合需要精细音频分析的场景,如智能监控、内容理解

近期多模态音视频语言模型在文本-音频检索方面表现优异,但在帧级音频理解上仍存在不足。以往方法虽使用时序感知标签或无监督训练提升帧级能力,但缺乏对事件发生时刻的细粒度标注。传统声音事件检测模型虽能精确定位事件,却仅限于预定义类别,在面对分布外事件时效果不佳。本文提出FLAM,一种开放词汇的对比音频-语言模型,可实现特定声音事件的定位。FLAM采用高效内存与校准的帧级目标函数,并通过逻辑调整缓解训练中的虚假相关性(如事件依赖、标签不平衡)。为实现帧级监督,我们利用包含多样化音频事件的大规模数据集,结合大语言模型生成的描述和仿真数据。实验与案例研究显示,FLAM显著提升了开放词汇的声音事件定位能力,同时保持了出色的全局检索与下游任务性能。

原文摘要 · Abstract (English)

Recent multi-modal audio-language models (ALMs) excel at text-audio retrieval but struggle with frame-wise audio understanding. Prior works use temporal-aware labels or unsupervised training to improve frame-wise capabilities, but they still lack fine-grained labeling capability to pinpoint when an event occurs. While traditional sound event detection models can precisely localize events, they are limited to pre-defined categories, making them ineffective for real-world scenarios with out-of-distribution events. In this work, we introduce FLAM, an open-vocabulary contrastive audio-language model capable of localizing specific sound events. FLAM employs a memory-efficient and calibrated frame-wise objective with logit adjustment to address spurious correlations, such as event dependencies and label imbalances during training. To enable frame-wise supervision, we leverage a large-scale dataset with diverse audio events, LLM-generated captions and simulation. Experimental results and case studies demonstrate that FLAM significantly improves the open-vocabulary localization capability while maintaining strong performance in global retrieval and downstream tasks.

音频理解多模态开放词汇事件定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。