开创视频流畅度评估新任务,构建首个专用数据集与基准
Pioneering Perceptual Video Fluency Assessment: A Novel Task with Benchmark Dataset and Baseline
- 提出独立于画质评估的视频流畅度评估任务
- 构建4606段真实场景视频数据集FluVid,含标准化评分体系
- 提出FluNet模型,利用时序置换自注意力提升长程帧间关联
准确估计人类对视频流畅度(如运动连贯性、帧间连续性)的主观反馈,在流媒体和游戏等领域至关重要。然而长期以来被忽视,以往工作将其作为视频质量评估(VQA)的一个子维度处理。本文通过初步实验发现,现有VQA方法对流畅度的预测严重不足,限制了应用效果。为此,我们首次将视频流畅度评估(VFA)作为一个独立的感知任务提出,面向时间维度。为推动该研究:1)构建了聚焦流畅度的数据集FluVid,包含4,606段真实视频,具有均衡的流畅度分布,是首个针对VFA设计的评分标准与人工评测;2)建立了涵盖23种方法的大型基准,为定制化模型设计提供洞见;3)提出基线模型FluNet,采用时序置换自注意力(T-PSA)增强输入流畅度信息,强化长距离帧间交互。本工作不仅达到当前最优性能,更向社区提供探索VFA解决方案的路线图。
原文摘要 · Abstract (English)
Accurately estimating humans' subjective feedback on video fluency, e.g., motion consistency and frame continuity, is crucial for various applications like streaming and gaming. Yet, it has long been overlooked, as prior arts have focused on solving it in the video quality assessment (VQA) task, merely as a sub-dimension of overall quality. In this work, we conduct pilot experiments and reveal that current VQA predictions largely underrepresent fluency, thereby limiting their applicability. To this end, we pioneer Video Fluency Assessment (VFA) as a standalone perceptual task focused on the temporal dimension. To advance VFA research, 1) we construct a fluency-oriented dataset, FluVid, comprising 4,606 in-the-wild videos with balanced fluency distribution, featuring the first-ever scoring criteria and human study for VFA. 2) We develop a large-scale benchmark of 23 methods, the most comprehensive one thus far on FluVid, gathering insights for VFA-tailored model designs. 3) We propose a baseline model called FluNet, which deploys temporal permuted self-attention (T-PSA) to enrich input fluency information and enhance long-range inter-frame interactions. Our work not only achieves state-of-the-art performance but, more importantly, offers the community a roadmap to explore solutions for VFA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。