arXiv:2505.13062cs.MMcs.SD2025-05中稿 · Interspeech 2025被引 1

让视觉语言模型从无声视频推断声音,突破跨模态推理瓶颈

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model

  • 基于思维链的微调策略提升模型跨模态推理能力
  • 在无声视频上生成音频描述准确率显著提升
  • 适合研究多模态理解与生成的学者参考

人类能从无声视频中直观推断出声音,但多模态大语言模型在不访问目标模态的情况下进行跨模态推理的能力仍待探索。现有文本辅助视频到音频(VT2A)方法在视频拟音任务中表现良好,但在推理阶段难以获取音频描述。本文提出从无声视频推理音频描述(SVAD)新任务,并研究视觉语言模型(VLMs)在此任务上的能力。为增强VLMs对SVAD任务的推理能力,我们构建了CoT-AudioCaps数据集,并提出基于思维链的监督微调策略。在SVAD及后续VT2A任务上的实验表明,该方法在两个关键方面均有效:显著提升VLMs对无声视频的跨模态推理能力,并成功解决VT2A推理中音频描述获取难题。

原文摘要 · Abstract (English)

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current text-assisted-video-to-audio (VT2A) methods excel in video foley tasks but struggle to acquire audio descriptions during inference. We introduce the task of Reasoning Audio Descriptions from Silent Videos (SVAD) to address this challenge and investigate vision-language models' (VLMs) capabilities on this task. To further enhance the VLMs' reasoning capacity for the SVAD task, we construct a CoT-AudioCaps dataset and propose a Chain-of-Thought-based supervised fine-tuning strategy. Experiments on SVAD and subsequent VT2A tasks demonstrate our method's effectiveness in two key aspects: significantly improving VLMs' modal-mismatch reasoning for SVAD and effectively addressing the challenge of acquiring audio descriptions during VT2A inference.

跨模态推理视觉语言模型音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。