让声音与文字在时空上精准对齐,提升多事件音频理解能力
CoSTALA: Compositional Spatio-Temporal Audio-Language Alignment via Multi-Grain Hierarchical Contrastive Learning
- 采用分层对比学习,从全局到细粒度建模声音与文本的时空关系
- 能准确区分多个连续声音事件,保持每个事件的语义完整性
- 适合需要精细音频场景理解的应用,如智能助手、视频字幕生成
传统音频语言模型在听觉与文本表征对齐方面取得显著进展,包括空间音频的探索。然而,在日常空间场景中,仍难以有效处理多事件音频序列。现有方法主要依赖全局听觉与文本特征的粗粒度对比学习,缺乏区分多个顺序事件的能力。为此,我们提出CoSTALA——一种从纯全局对齐转向细粒度时空推理的新训练范式。通过构建多粒度分层损失函数体系,实现对时间依赖性的显式建模,并成功将各声学事件锚定,保持其语义纯净性。大量实验表明,CoSTALA显著建立了一个强大的新型时空音频理解框架。
原文摘要 · Abstract (English)
Conventional audio language models (ALMs) have made significant progress in achieving alignment between auditory and textual representations, including recent explorations in spatial audio. However, in daily spatial scenarios, they still cannot effectively process multi-event audio sequences. Current approaches primarily rely on coarse-grained contrastive learning with global auditory and textual features, lacking the resolution to distinguish multiple sequential events. To overcome these limitations, we propose CoSTALA-a novel training paradigm that transitions from purely global alignment to fine-grained spatio-temporal reasoning. By constructing a multi-granularity hierarchical loss function system, we achieve explicit modeling of temporal dependencies, and successfully anchors individual acoustic events to preserve their semantic purity. Extensive experiments demonstrate that CoSTALA significantly establish a powerful new framework for spatio-temporal audio understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。