剖析文本到音频检索模型对声音时序的理解能力
Dissecting Temporal Understanding in Text-to-Audio Retrieval
- 分析主流模型在音频时序理解上的表现
- 提出新数据集,可精准评估时序建模能力
- 适合关注音频生成与理解的科研人员
近年来,机器学习的发展推动了多模态任务的研究,如文本到视频和文本到音频检索。这些任务要求模型理解视频和音频数据的语义内容,包括物体和角色,并学习空间布局与时间关系。本文聚焦于声音时序顺序这一被忽视的问题,在AudioCaps和Clotho数据集上剖析了一个先进的文本到音频检索模型的时序理解能力。此外,我们构建了一个合成的文本-音频数据集,为评估近期模型的时序能力提供受控环境。最后,提出一种损失函数,促使模型关注事件的时间顺序。代码与数据已公开。
原文摘要 · Abstract (English)
Recent advancements in machine learning have fueled research on multimodal tasks, such as for instance text-to-video and text-to-audio retrieval. These tasks require models to understand the semantic content of video and audio data, including objects, and characters. The models also need to learn spatial arrangements and temporal relationships. In this work, we analyse the temporal ordering of sounds, which is an understudied problem in the context of text-to-audio retrieval. In particular, we dissect the temporal understanding capabilities of a state-of-the-art model for text-to-audio retrieval on the AudioCaps and Clotho datasets. Additionally, we introduce a synthetic text-audio dataset that provides a controlled setting for evaluating temporal capabilities of recent models. Lastly, we present a loss function that encourages text-audio models to focus on the temporal ordering of events. Code and data are available at https://www.robots.ox.ac.uk/~vgg/research/audio-retrieval/dtu/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。