自动挑选或生成高质量视频缩略图,提升效率与吸引力。
Automating Video Thumbnails Selection and Generation with Multimodal and Multistage Analysis
- 分多阶段处理,融合视觉与语义分析选图或生成新图。
- 82人调查中45.77%偏好该方法,专业设计候选数提升3.57倍。
- 适合需要批量生成高质缩略图的视频平台与内容创作者。
本论文提出一种自动化视频缩略图选择与生成方法,适用于传统广播内容。方法设定严格标准,确保缩略图多样性、代表性与美学质量,考虑标志位置空间、垂直画幅适配性,以及人脸身份与情绪的准确识别。提出多阶段流水线,可从视频中选取候选帧,或通过融合视频元素、使用扩散模型生成新图像。流水线整合了最先进的模型,涵盖下采样、冗余减少、自动裁剪、人脸检测、闭眼与情绪识别、镜头尺度与审美预测、分割、抠图及色调统一等任务,并利用大语言模型与视觉变换器保证语义一致性。配套图形界面工具支持快速浏览输出结果。为评估方法,进行了全面实验:在69个视频的测试中,53.6%的推荐集包含专业设计师选定的缩略图,其中73.9%包含相似图像;82名参与者调查显示,45.77%更偏好该方法,高于手动选图的37.99%和替代方法的16.36%。专业设计师表示,有效候选数量较替代方法提升3.57倍,验证了方法满足既定标准。结论表明,该方法在加速缩略图制作的同时保持高质量,有助于提升用户参与度。
原文摘要 · Abstract (English)
This thesis presents an innovative approach to automate video thumbnail selection for traditional broadcast content. Our methodology establishes stringent criteria for diverse, representative, and aesthetically pleasing thumbnails, considering factors like logo placement space, incorporation of vertical aspect ratios, and accurate recognition of facial identities and emotions. We introduce a sophisticated multistage pipeline that can select candidate frames or generate novel images by blending video elements or using diffusion models. The pipeline incorporates state-of-the-art models for various tasks, including downsampling, redundancy reduction, automated cropping, face recognition, closed-eye and emotion detection, shot scale and aesthetic prediction, segmentation, matting, and harmonization. It also leverages large language models and visual transformers for semantic consistency. A GUI tool facilitates rapid navigation of the pipeline's output. To evaluate our method, we conducted comprehensive experiments. In a study of 69 videos, 53.6% of our proposed sets included thumbnails chosen by professional designers, with 73.9% containing similar images. A survey of 82 participants showed a 45.77% preference for our method, compared to 37.99% for manually chosen thumbnails and 16.36% for an alternative method. Professional designers reported a 3.57-fold increase in valid candidates compared to the alternative method, confirming that our approach meets established criteria. In conclusion, our findings affirm that the proposed method accelerates thumbnail creation while maintaining high-quality standards and fostering greater user engagement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。