arXiv:2604.14816cs.CVcs.HC2026-04被引 22

2000段视频+5000人眼动数据,挑战自动预测视觉注意力点

NTIRE 2026 Challenge on Video Saliency Prediction: Methods and Results

论文配图:NTIRE 2026 Challenge on Video Saliency Prediction: Methods and Results
图 1 · 摘自论文原文
  • 用众包鼠标轨迹收集真实观看行为,构建大规模视频显著性数据集
  • 7支队伍通过代码审查,最佳模型在800个测试视频上表现最优
  • 适合关注视频注意力建模、人眼追踪与数据集构建的研究者

本文介绍了NTIRE 2026视频显著性预测挑战赛的背景与成果。挑战旨在推动自动视频显著性图预测方法的发展。主办方构建了一个包含2000段多样化视频的公开数据集,采用众包鼠标追踪技术采集超过5000名评估者的注视点与对应显著性图。评测在800个测试视频上使用主流质量指标进行。共有20多支团队参与提交,其中7支队伍通过最终代码审查。所有数据与代码已公开:https://github.com/msu-video-group/NTIRE26_Saliency_Prediction。

原文摘要 · Abstract (English)

This paper presents an overview of the NTIRE 2026 Challenge on Video Saliency Prediction. The goal of the challenge participants was to develop automatic saliency map prediction methods for the provided video sequences. The novel dataset of 2,000 diverse videos with an open license was prepared for this challenge. The fixations and corresponding saliency maps were collected using crowdsourced mouse tracking and contain viewing data from over 5,000 assessors. Evaluation was performed on a subset of 800 test videos using generally accepted quality metrics. The challenge attracted over 20 teams making submissions, and 7 teams passed the final phase with code review. All data used in this challenge is made publicly available - https://github.com/msu-video-group/NTIRE26_Saliency_Prediction.

视频显著性众包数据人眼追踪竞赛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。