arXiv:2412.00049cs.MMcs.AI2024-12综述被引 15

综述音频视觉关联学习最新进展与挑战

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

  • 系统梳理音视频关联学习的模型与方法
  • 分析主流目标函数与跨模态表征策略
  • 适合多媒体AI研究者参考

音频-视觉关联学习旨在捕捉和理解音频与视觉数据间的自然关联。深度学习的快速发展推动了大量处理音视频数据的方法涌现,促使开展全面综述。本文不仅分析该领域所用模型,还探讨任务定义、学习范式及常用目标函数,并研究音频-视觉数据在优化过程中的利用方式,即不同知识表示方法。重点在于如何通过人类可理解的机制——即反映可解释知识的结构化知识——引导学习过程。最后,总结音视频关联学习(AVCL)的最新进展,并讨论未来研究方向。

原文摘要 · Abstract (English)

Audio-visual correlation learning aims to capture and understand natural phenomena between audio and visual data. The rapid growth of Deep Learning propelled the development of proposals that process audio-visual data and can be observed in the number of proposals in the past years. Thus encouraging the development of a comprehensive survey. Besides analyzing the models used in this context, we also discuss some tasks of definition and paradigm applied in AI multimedia. In addition, we investigate objective functions frequently used and discuss how audio-visual data is exploited in the optimization process, i.e., the different methodologies for representing knowledge in the audio-visual domain. In fact, we focus on how human-understandable mechanisms, i.e., structured knowledge that reflects comprehensible knowledge, can guide the learning process. Most importantly, we provide a summarization of the recent progress of Audio-Visual Correlation Learning (AVCL) and discuss the future research directions.

音视频关联多模态学习综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。