arXiv:2605.29302cs.CV2026-05

用深度模型预测短视频广告的注意力分布,帮广告设计更快更高效。

ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement

论文配图:ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement
图 1 · 摘自论文原文
  • 基于3D U-Net架构,融合音效与场景语义信息预测动态注意力图。
  • 在151条广告、每条20人眼动数据上测试,显著提升预测准确率。
  • 通过熵值分析识别低吸引力片段,适合广告优化与A/B测试使用。

数字媒体环境正加速向短时视频广告转型,涵盖电视、社交媒体及电商平台。本研究聚焦短形式视频广告的深层注意力预测。现有深度注意力模型可预测人类注视轨迹,以提升人机交互体验并优化设计。对于视频广告,动态注意力图能揭示观众何时何地关注内容,从而解释广告有效性并指导优化。本文提出并验证了一种新型深度动态注意力预测模型——ViASNet(Video Ad Saliency Network),其架构基于3D U-Net,融合音频与场景语义信息。模型在151条视频广告上进行评估,每条广告有约20名受试者的眼动追踪数据,并通过消融实验分析关键影响因素。我们逐帧计算预测注意力图的熵值,作为诊断工具识别未吸引观众的广告与场景,并在15条未见广告数据上验证其应用效果。研究表明,基于此类深度注意力模型(如ViASNet)构建的自动化系统可大幅加速广告设计与测试流程。

原文摘要 · Abstract (English)

The digital media landscape has seen a pervasive shift toward short-form video advertising on TV, social media and e-commerce platforms. The present study focuses on deep saliency prediction for short-form video advertising. Deep saliency models have been used to generate predictions of human eye fixation patterns with the purpose of enhancing user interaction with digital technology and optimizing its design. For video ads, dynamic saliency maps capture where and when viewers are looking, revealing why video ads are effective, and how their content should be optimized. We develop and test a new deep dynamic saliency prediction model called ViASNet (Video Ad Saliency Network), which has an architecture founded on the 3D U-Net, and accommodates the influence of audio and the semantic meaning of scenes. We assess the model's performance on 151 video ads, each seen by about 20 viewers wile their eye movements were tracked, and explore the critical factors influencing model performance through ablation experiments. We calculate the entropy of the predicted saliency maps frame-by-frame as a diagnostic tool to identify ads and scenes that fail to engage viewers, and illustrate its use on test data of 15 unseen ads. Our study reveals that ad design and testing can be sped up considerably through automated systems built on deep saliency models such as ViASNet.

视频广告注意力预测3D U-Net眼动分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。