用智能流程自动对齐音视频数据,提升多模态表示质量
Aligning Audio-Visual Joint Representations with an Agentic Workflow
- 通过大模型驱动的智能流程,分步完成音视频描述、判断对齐状态、编辑音频
- 在多个下游任务中实现当前最佳性能,显著提升音视频联合表示效果
- 适合需要高质量音视频对齐的研究者或开发者,尤其关注数据质量提升
视觉内容与伴随的音频信号天然构成联合表征,有助于提升音视频相关应用。尽管已有多种音视频表征学习框架,但音视频数据对齐的重要性常被忽视,影响表征质量。本文观察到音频可能含背景噪声,且音视频流之间可能存在不同步问题,这些非严格对齐限制了表征质量并降低应用性能。为此,我们提出从数据中心视角改进音视频联合表征,通过一个由大语言模型驱动的助手 AVAgent 实现音频信号与视觉数据的对齐。对于每一对输入音视频数据,AVAgent 使用多模态大模型分别将音频和视觉数据转化为语言描述(工具使用);随后推理该配对是否对齐良好,并规划是否需编辑音频(规划);音频编辑通过预定义操作执行,如降噪或数据增强;此外,使用视觉语言模型评估修改后的音频与视觉内容的匹配度,并提供反馈给 AVAgent(反思)。工具使用、规划与反思步骤循环进行,形成智能工作流,使音频信号逐步对齐视觉内容。现有方法可直接利用该流程生成的对齐音视频数据,以提升音视频联合表征。实验结果全面证明,所提方法在多样下游任务中优于以往基线,达到当前最优表现。
原文摘要 · Abstract (English)
Visual content and accompanied audio signals naturally formulate a joint representation to improve audio-visual (AV) related applications. While studies develop various AV representation learning frameworks, the importance of AV data alignment is usually undermined for achieving high-quality representation. We observe that an audio signal may contain background noise interference. Also, non-synchronization may appear between audio and video streams. These non-strict data alignment limits representation quality and downgrade application performance. In this paper, we propose to improve AV joint representations from a data-centric perspective by aligning audio signals to visual data. Our alignment is conducted in an agentic workflow controlled by an LLM-based assistant named AVAgent. For each input AV data pair, our AVAgent uses a multi-modal LLM to convert audio and visual data into language descriptions separately (i.e., tool use). Then, AVAgent reasons whether this paired data is aligned well and plans to edit the audio signal if needed (i.e., planning). The audio editing is executed by predefined actions that filter noise or augment data. Moreover, we use a VLM to evaluate how modified audio signals match the visual content and provide feedback to AVAgent (i.e., reflection). The tool use, planning, and reflection steps operate cyclically to become an agentic workflow where audio signals are gradually aligned to visual content. To this end, existing methods can directly leverage the aligned AV data via our agentic workflow to improve AV joint representations. The experimental results comprehensively demonstrate the state-of-the-art performance of the proposed approach against previous baselines in diverse downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。