arXiv:2504.14267cs.CV2025-04被引 4

用图文音三模态扩散模型预测视频显著性,精度提升明显。

Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction

  • 将显著性预测转为图文音条件下的图像生成任务,通过逐步去噪生成结果。
  • 在SIM、CC、NSS、AUC-J上分别提升1.03%、2.35%、2.71%、0.33%。
  • 设计SITR机制与Saliency-DiT结构,有效融合多模态信息,适合多模态视觉研究者。

视频显著性预测对视频压缩和人机交互等下游任务至关重要。随着多模态学习的发展,研究者开始探索音频-视觉和文本-视觉的多模态显著性预测。听觉线索引导观众视线聚焦声源,文本线索提供语义指导以理解视频内容。融合这些互补线索可提升预测精度。本文同时分析视觉、听觉和文本模态,提出TAVDiff:一种文本-音频-视觉条件扩散模型用于视频显著性预测。TAVDiff将显著性预测视为基于文本、音频和视觉输入的图像生成任务,通过逐步去噪生成显著图。为有效利用文本,采用大型多模态模型生成视频帧的文本描述,并引入面向显著性的图像-文本响应(SITR)机制生成图像-文本响应图,作为条件信息引导模型定位与文本描述语义相关的视觉区域。听觉模态则作为另一条件信息,引导模型关注由声音指示的显著区域。同时,由于扩散变换器(DiT)直接将条件信息与时间步拼接,可能影响噪声水平估计。为此,我们提出Saliency-DiT,将条件信息与时间步解耦。实验表明,TAVDiff在SIM、CC、NSS、AUC-J指标上分别优于现有方法1.03%、2.35%、2.71%、0.33%。

原文摘要 · Abstract (English)

Video saliency prediction is crucial for downstream applications, such as video compression and human-computer interaction. With the flourishing of multimodal learning, researchers started to explore multimodal video saliency prediction, including audio-visual and text-visual approaches. Auditory cues guide the gaze of viewers to sound sources, while textual cues provide semantic guidance for understanding video content. Integrating these complementary cues can improve the accuracy of saliency prediction. Therefore, we attempt to simultaneously analyze visual, auditory, and textual modalities in this paper, and propose TAVDiff, a Text-Audio-Visual-conditioned Diffusion Model for video saliency prediction. TAVDiff treats video saliency prediction as an image generation task conditioned on textual, audio, and visual inputs, and predicts saliency maps through stepwise denoising. To effectively utilize text, a large multimodal model is used to generate textual descriptions for video frames and introduce a saliency-oriented image-text response (SITR) mechanism to generate image-text response maps. It is used as conditional information to guide the model to localize the visual regions that are semantically related to the textual description. Regarding the auditory modality, it is used as another conditional information for directing the model to focus on salient regions indicated by sounds. At the same time, since the diffusion transformer (DiT) directly concatenates the conditional information with the timestep, which may affect the estimation of the noise level. To achieve effective conditional guidance, we propose Saliency-DiT, which decouples the conditional information from the timestep. Experimental results show that TAVDiff outperforms existing methods, improving 1.03\%, 2.35\%, 2.71\% and 0.33\% on SIM, CC, NSS and AUC-J metrics, respectively.

视频显著性扩散模型多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。