让视觉模型通过看懂艺术中的情感,仅用少量音频数据就能听懂
Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound
- 用视觉引导音频对齐,无需大规模音频预训练即可学会听
- 在艺术情感理解上超越单模态与多模态基线方法
- 适合研究跨模态情感理解、艺术与AI融合的学者
情感理解对提升大语言模型的通用性、可靠性及与人类的一致性至关重要。艺术通过视觉与听觉元素的协同设计传递情感,但以往研究多以人为中心或仅处理单一模态,忽视了作品本身表达的情感意图。当前音频-视觉语言模型通常需大规模音频预训练才能赋予视觉模型听觉能力,限制了可扩展性。我们提出视觉锚定的音视频情感大模型(VAEmotionLLM),采用两阶段框架:第一阶段,视觉引导音频对齐(VG-Align)通过同步音视频片段中共享大模型的下一步词分布对齐,将冻结的视觉路径知识蒸馏至音频路径,实现无需大规模音频数据的听觉能力;第二阶段,轻量级跨模态情感适配器(EmoAdapter)包含情感增强器和情感监督器,注入情感敏感残差并施加情感监督,提升跨模态情感理解。我们还构建了艺术中心的情感评估基准ArtEmoBenchmark,评估在纯音频、纯视觉及音视频联合输入下的内容与情感理解能力。VAEmotionLLM在ArtEmoBenchmark上达到当前最优性能,优于各类基线模型。消融实验表明各组件具有互补性。
原文摘要 · Abstract (English)
Emotion understanding is critical for making Large Language Models (LLMs) more general, reliable, and aligned with humans. Art conveys emotion through the joint design of visual and auditory elements, yet most prior work is human-centered or single-modality, overlooking the emotion intentionally expressed by the artwork. Meanwhile, current Audio-Visual Language Models (AVLMs) typically require large-scale audio pretraining to endow Visual Language Models (VLMs) with hearing, which limits scalability. We present Vision Anchored Audio-Visual Emotion LLM (VAEmotionLLM), a two-stage framework that teaches a VLM to hear by seeing with limited audio pretraining and to understand emotion across modalities. In Stage 1, Vision-Guided Audio Alignment (VG-Align) distills the frozen visual pathway into a new audio pathway by aligning next-token distributions of the shared LLM on synchronized audio-video clips, enabling hearing without a large audio dataset. In Stage 2, a lightweight Cross-Modal Emotion Adapter (EmoAdapter), composed of the Emotion Enhancer and the Emotion Supervisor, injects emotion-sensitive residuals and applies emotion supervision to enhance cross-modal emotion understanding. We also construct ArtEmoBenchmark, an art-centric emotion benchmark that evaluates content and emotion understanding under audio-only, visual-only, and audio-visual inputs. VAEmotionLLM achieves state-of-the-art results on ArtEmoBenchmark, outperforming audio-only, visual-only, and audio-visual baselines. Ablations show that the proposed components are complementary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。