arXiv:2606.20418cs.SD2026-06中稿 · Interspeech 2026

用混合音频对建模声音重叠,提升音文对齐的不确定性表达

MixProLAP: Mixture-Induced Uncertainty Modeling for Probabilistic Language-Audio Pretraining

  • 通过混合音文对生成真实混响场景,模拟多对多对应关系
  • 在多个音文检索任务上优于传统确定性方法,提升对齐精度
  • 适合研究音视频多模态融合与不确定性建模的学者

声学环境常包含多个重叠的声音事件,同一声景也可用不同文本描述,导致音文对齐本质上存在歧义。本文提出一种概率化音文预训练框架,用于建模音文对齐中的多对多对应模糊性。不同于传统对比学习中使用确定性点嵌入的方法,本方法将每种模态表示为分布,并学习具备不确定性感知的跨模态对齐。不依赖掩码模拟不确定性,而是通过混合音文对生成重叠声音,更贴近真实声学混合场景,并捕捉声音事件间的语义包含关系。进一步引入多层次包含损失,强制表示符合这些关系。在多个音文检索基准上的实验表明,所提方法显著优于确定性基线。

原文摘要 · Abstract (English)

Acoustic environments often contain multiple overlapping sound events, and the same acoustic scene can be described using diverse textual expressions, making audio-text alignment inherently ambiguous. This paper proposes a probabilistic audio-language pretraining framework to model many-to-many correspondence ambiguity in audio-text alignment. Unlike conventional contrastive methods that learn deterministic point embeddings, our approach represents each modality as a distribution and learns uncertainty-aware cross-modal alignment. Rather than relying on masking-based uncertainty simulation, we mix audio-text pairs to create overlapping sounds that better reflect real acoustic mixtures and capture semantic inclusion relations among sound events. We further introduce a multi-level inclusion loss to enforce representations consistent with these relations. Experiments on audio-text retrieval benchmarks show that the proposed method outperforms deterministic baselines.

音文对齐不确定性建模多模态预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。