用生成模型特征提升音频理解,兼顾细节感知与语义认知
When Audio Generators Become Good Listeners: Generative Features for Understanding Tasks
- 利用生成模型提取兼具时频细节与语义先验的音频特征
- 在声音事件分类、标签识别和细粒度音频描述任务中均实现性能提升
- 为音频表征学习提供生成与判别互补的新视角,适合音频理解研究者
本工作首次探索将生成式特征用于增强音频理解。与传统判别式特征直接优化后验概率、侧重语义抽象而丢失细节不同,音频生成模型天然同时编码时空感知(捕捉时间与频率上的局部声学纹理)和语义先验(知道该生成什么)。这促使我们探究二者互补性的桥梁。本文系统分析了两类特征的差异与互补关系,并提出有效的融合策略。在多个任务上——包括声音事件分类、标签识别,以及细粒度的音频描述任务——实验均显示稳定性能提升。除了实证效果,本工作更重要的是引入了音频表示学习的新视角:生成与判别互补性可为音频理解同时提供精细感知与语义意识。
原文摘要 · Abstract (English)
This work pioneers the utilization of generative features in enhancing audio understanding. Unlike conventional discriminative features that directly optimize posterior and thus emphasize semantic abstraction while losing fine grained details, audio generation models inherently encode both spatiotemporal perception (capturing local acoustic texture across time and frequency) and semantic prior (knowing what to generate). It motivates us to explore the bridge of these complementary strengths. We provide a systematic investigation of their differences and complementary relationships, and ultimately propose an effective fusion strategy. Experiments across multiple tasks, including sound event classification, tagging, and particularly the fine grained task of audio captioning, demonstrate consistent performance gains. Beyond empirical improvements, this work more importantly introduces a new perspective on audio representation learning, highlighting that generative discriminative complementarity can provide both detailed perception and semantic awareness for audio understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。