arXiv:2605.20838cs.CVcs.AI2026-05

构建首个用户生成短视频数据集,推动视频高层语义理解研究

USV: Towards Understanding the User-generated Short-form Videos

论文配图:USV: Towards Understanding the User-generated Short-form Videos
图 1 · 摘自论文原文
  • 从UGC平台采集22.4万条短视频,无人工筛选
  • 提出主题识别与图文检索任务,验证模型性能
  • 适合研究短视频理解、多模态学习的学者和工程师

近年来,多个大规模视频数据集推动了视频理解的发展。然而,新兴的用户生成短时视频仍缺乏系统研究。本文提出USV数据集,用于高层语义视频理解。该数据集包含约22.4万条来自UGC平台的视频,通过标签查询获取,未经过额外人工验证或剪辑。尽管视频理解近年取得显著进展,但多数工作聚焦实例级识别,难以捕捉视频的高层语义信息。为此,我们进一步构建两个任务:主题识别与视频-文本检索。提出两种统一高效的基线方法:多模态融合网络(MMF-Net)用于主题识别,视频-文本对比学习(VTCL)用于视频-文本检索,并开展全面基准测试,以促进后续研究。项目主页:https://usvdataset.github.io。

原文摘要 · Abstract (English)

Several large-scale video datasets have been published these years and have advanced the area of video understanding. However, the newly emerged user-generated short-form videos have rarely been studied. This paper presents USV, the User-generated Short-form Video dataset for high-level semantic video understanding. The dataset contains around 224K videos collected from UGC platforms by label queries without extra manual verification and trimming. Although video understanding has achieved plausible improvement these years, most works focus on instance-level recognition, which is not sufficient for learning the representation of the high-level semantic information of videos. Therefore, we further establish two tasks: topic recognition and video-text retrieval on USV. We propose two unified and effective baseline methods Multi-Modality Fusion Network (MMF-Net) and Video-Text Contrastive Learning (VTCL), to tackle the topic recognition task and video-text retrieval respectively, and carry out comprehensive benchmarks to facilitate future research. Our project page is https://usvdataset.github.io.

视频理解短视频多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。