构建首个大规模视频美学数据库,支持多维度专业评分。
VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations
- 采集10,490段视频,由37位专业人士标注多维美学分数。
- 提出VADB-Net模型,在美学评分任务上超越现有方法。
- 适合视频内容评估、审美生成等研究者使用。
视频美学评估是多媒体计算中的关键领域,融合计算机视觉与人类认知。其发展受限于缺乏标准化数据集和鲁棒模型,因视频的时间动态性和多模态融合特性难以直接应用图像评估方法。本研究推出VADB,目前最大的视频美学数据库,包含10,490段多样视频,由37位专业人士在整体与属性特定美学得分、丰富语言评论及客观标签等多个维度进行标注。我们提出VADB-Net,一种双模态预训练框架,采用两阶段训练策略,在评分任务中优于现有视频质量评估模型,并可支持下游视频美学评估任务。数据集与源代码已开源:https://github.com/BestiVictory/VADB。
原文摘要 · Abstract (English)
Video aesthetic assessment, a vital area in multimedia computing, integrates computer vision with human cognition. Its progress is limited by the lack of standardized datasets and robust models, as the temporal dynamics of video and multimodal fusion challenges hinder direct application of image-based methods. This study introduces VADB, the largest video aesthetic database with 10,490 diverse videos annotated by 37 professionals across multiple aesthetic dimensions, including overall and attribute-specific aesthetic scores, rich language comments and objective tags. We propose VADB-Net, a dual-modal pre-training framework with a two-stage training strategy, which outperforms existing video quality assessment models in scoring tasks and supports downstream video aesthetic assessment tasks. The dataset and source code are available at https://github.com/BestiVictory/VADB.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。