用概率分布建模音视频多对多关系,提升语义理解能力
ProLAP: Probabilistic Language-Audio Pre-Training
- 将音视频映射到概率分布空间,捕捉多对多语义关联
- 在小数据下仍能学习层次化语义结构,优于以往方法
- 适合音视频检索与语义遍历任务,尤其关注语义多样性
语言-音频联合表征学习框架通常依赖确定性嵌入,假设音频与文本间存在一一对应关系。但在真实场景中,这种关系本质上是多对多的:一段音频可被多个描述词句表达,反之亦然。为此,我们提出概率语言-音频预训练(ProLAP),将多重性建模为联合语言-音频嵌入空间中的概率分布扩散。为有效训练模态内层次关系,引入两个新目标:(i) 层次包含损失,促进输入的语义层次理解;(ii) 掩码排斥损失,提升优化层次包含损失时的学习效率。该训练策略使模型即使在小数据上也能学习数据固有的层次结构,而无需大规模数据支持。实验表明,ProLAP在音频-文本检索任务中优于现有确定性方法。此外,通过本文提出的音频遍历任务实验,验证了其能捕捉合理的语义层次结构。
原文摘要 · Abstract (English)
Language-audio joint representation learning frameworks typically depend on deterministic embeddings, assuming a one-to-one correspondence between audio and text. In real-world settings, however, the language-audio relationship is inherently many-to-many: one audio segment can be described by multiple captions and vice versa. To address this, we propose Probabilistic Language-Audio Pre-training (ProLAP), which models multiplicity as the spread of probability distributions in a joint language-audio embedding space. To train the intra-modal hierarchical relationship effectively, we also introduce two objectives: (i) hierarchical inclusion loss to promote semantic hierarchical understanding of inputs and (ii) mask repulsive loss to improve the efficiency of learning when optimizing the hierarchical inclusion loss. With this training strategy, our model can learn the hierarchical structure inherent in the data even from small datasets, in contrast to prior probabilistic approaches that rely on large-scale datasets. In our experiments, ProLAP outperforms existing deterministic approaches on audio-text retrieval tasks. Moreover, through experiments on the audio traversal task introduced in this paper, we demonstrate that ProLAP captures the plausible semantic hierarchy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。