arXiv:2504.11733cs.CV2025-04被引 8

提出DVLTA-VQA模型,用文本引导融合视频质量评估中的视觉与语言信息。

DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment

  • 分离CLIP的视觉与语言模块,分别对应人眼的物体识别与运动感知路径。
  • 引入时序模块显式建模视频动态,提升对运动信息的捕捉能力。
  • 适合关注视频质量评估中多模态融合与动态特征整合的研究者。

受人类视觉系统双通路理论启发——腹侧通路负责物体识别与细节分析,背侧通路关注空间关系与运动感知——越来越多视频质量评估(VQA)方法基于此框架提出。近年来,大规模多模态模型如对比语言-图像预训练(CLIP)的进展,推动研究者将CLIP融入双通路型VQA方法中,以利用其强大的语义理解能力,模拟腹侧通路的物体识别与细节分析,以及背侧通路的空间关系分析。然而,CLIP原生针对图像设计,缺乏对视频固有时序与运动信息的建模能力。为解决此问题,本文提出一种用于无参考视频质量评估(NR-VQA)的解耦视觉-语言建模与文本引导自适应方法(DVLTA-VQA),将CLIP的视觉与文本组件解耦,并分别集成到NR-VQA流程的不同阶段。具体而言,提出视频时序CLIP模块以显式建模时序动态,增强运动感知,契合背侧通路;同时设计时序上下文模块以细化帧间依赖,进一步提升运动建模能力。在腹侧通路侧,采用基础视觉特征提取模块强化细节分析。最后,提出文本引导的自适应融合策略,实现特征的动态加权,促进空间与时序信息的有效整合。

原文摘要 · Abstract (English)

Inspired by the dual-stream theory of the human visual system (HVS) - where the ventral stream is responsible for object recognition and detail analysis, while the dorsal stream focuses on spatial relationships and motion perception - an increasing number of video quality assessment (VQA) works built upon this framework are proposed. Recent advancements in large multi-modal models, notably Contrastive Language-Image Pretraining (CLIP), have motivated researchers to incorporate CLIP into dual-stream-based VQA methods. This integration aims to harness the model's superior semantic understanding capabilities to replicate the object recognition and detail analysis in ventral stream, as well as spatial relationship analysis in dorsal stream. However, CLIP is originally designed for images and lacks the ability to capture temporal and motion information inherent in videos. To address the limitation, this paper propose a Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment (DVLTA-VQA), which decouples CLIP's visual and textual components, and integrates them into different stages of the NR-VQA pipeline. Specifically, a Video-Based Temporal CLIP module is proposed to explicitly model temporal dynamics and enhance motion perception, aligning with the dorsal stream. Additionally, a Temporal Context Module is developed to refine inter-frame dependencies, further improving motion modeling. On the ventral stream side, a Basic Visual Feature Extraction Module is employed to strengthen detail analysis. Finally, a text-guided adaptive fusion strategy is proposed to enable dynamic weighting of features, facilitating more effective integration of spatial and temporal information.

视频质量评估多模态建模文本引导双通路

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。