通过音频描述对齐提升通用音频表征的语义精度。
Semantic Refinement of Universal Audio Representations through Audio-Description Alignment

- 在预训练编码器基础上加入音频-描述对齐任务,增强高阶语义理解。
- 线性探针分类准确率提升4.66点,序列感知大模型在字幕生成中表现更优。
- 适用于需要精准语义理解的多领域音频任务,如语音、音乐与环境音分析。
通用音频表征需在保留声学细节的同时,使跨语音、音乐、环境音等领域的高层概念对不同容量下游模型都可访问。本文通过向BEST-RQ、重建和CTC的基础架构中添加音频-描述对齐,研究编码器的语义精炼。对比配对、乱序描述及正确配对三种轨迹,区分正确对应关系与额外对比目标的影响。所有端点冻结后,使用时序平均线性探针和序列感知大模型读出进行评估,检验精炼信息是否直接可用且对强模型仍有价值。在三个配对种子下,正确对齐使线性探针的领域平衡分类提升4.66点,序列感知大模型提升2.59点,各领域均获正向增益。线性探针的4.66点提升中,87%归因于正确配对;大模型则在字幕生成任务中展现最明显的优势。密集声学目标在两种读出方式下均带来互补增益。一个独立的24层延续模型在共享评测器上仍可媲美领先公开编码器,支持该方法在控制实验外的适用性。
原文摘要 · Abstract (English)
Universal audio representations must preserve acoustic detail while making high-level concepts accessible across speech, music, environmental sound, and downstream models of different capacities. We study semantic refinement of an acoustically pretrained encoder by adding audio-description alignment to a foundation of BEST-RQ, reconstruction, and CTC. We compare matched control, shuffled-description, and correctly paired trajectories to distinguish correct correspondence from an extra contrastive objective. Each endpoint is frozen and evaluated with a temporal-mean linear probe and a sequence-aware LLM readout, testing whether the refined information is directly accessible and remains useful to a stronger model. Across three paired seeds, correct alignment improves domain-balanced classification by 4.66 points with the linear probe and 2.59 points with the sequence-aware LLM, with positive changes in every domain. Correct pairing accounts for 87% of the linear-probe gain, while the LLM shows its clearest correspondence-specific benefit in captioning. Dense acoustic objectives provide complementary gains under both readouts. A separate 24-layer continuation remains competitive with leading public encoders under the shared evaluator, supporting the recipe beyond the controlled study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。