arXiv:2605.11572cs.CV2026-05

用文本做桥梁,让音视频模型高效对齐语义。

TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning

论文配图:TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning
图 1 · 摘自论文原文
  • 以文本为语义锚点,桥接音视频特征
  • 在多个数据集上达到顶尖性能
  • 适合追求高效微调的音视频研究者

音频-视觉理解需要有效对齐异构模态,但当时间对齐的音视频信号缺乏明确语义对应时,跨模态对应仍具挑战。本文提出使用文本作为语义锚点进行音视频表征学习。为此,我们构建了一个基于冻结音视频编码器的参数高效适配框架——文本桥接音视频适配器(TB-AVA),实现文本介导的音视频流交互。核心机制是门控语义调制(GSM),根据文本推断的语义相关性选择性调制特征通道。我们在AVE、AVS和AVVP等多个基准上评估该方法,结果表明所提框架在音视频学习中实现了参数高效微调的最先进性能,验证了文本作为有效语义锚点的潜力。

原文摘要 · Abstract (English)

Audio-visual understanding requires effective alignment between heterogeneous modalities, yet cross-modal correspondence remains challenging when temporally aligned audio and visual signals lack clear semantic correspondence. We propose to use text as a semantic anchor for audio-visual representation learning. To this end, we introduce a parameter-efficient adaptation framework built on frozen audio and visual encoders, centered on Text-Bridged Audio-Visual Adapter (TB-AVA), which enables text-mediated interaction between audio and visual streams. At the core of TB-AVA, Gated Semantic Modulation (GSM) selectively modulates feature channels based on text-inferred semantic relevance. We evaluate the proposed approach on multiple benchmarks, including AVE, AVS, and AVVP, where the proposed framework achieves state-of-the-art performance, demonstrating text as an effective semantic anchor for parameter-efficient fine-tuning (PEFT) in audio-visual learning.

音视频对齐参数高效文本桥接

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。