用奇幻动画等非常规视频训练模型,提升开放世界视觉学习能力。
What Can We Learn from Harry Potter? An Exploratory Study of Visual Representation Learning from Atypical Videos
- 用科幻、动画等非常规视频进行训练,增强模型泛化能力。
- 引入多样化的非常规数据后,OOD检测与零样本动作识别性能提升。
- 少量语义多样性的异常视频比大量常见视频更有效。
人类在面对开放世界中不常见概念时表现出卓越的泛化与发现能力,但现有研究多聚焦于封闭集中的典型数据,对开放世界新概念发现的研究仍不足。本文探索在训练过程中引入非常规视频(如科幻片、动画等)的影响。我们构建了一个包含多种异常类型视频的新数据集,并在此基础上进行表示学习。针对开放世界三大任务——分布外检测(OOD)、新类别发现(NCD)和零样本动作识别(ZSAR),实验表明,仅使用非常规视频进行简单训练即可在多种设置下持续提升性能。增加非常规样本的类别多样性能进一步提升OOD检测效果;在NCD任务中,使用较小但语义更丰富的非常规样本集优于更大但更典型的集合;在ZSAR中,非常规视频的语义多样性有助于模型更好泛化至未见动作类别。这些结果揭示了非常规视频在开放世界视觉学习中的价值,并提供了新数据集以推动后续研究。项目页面:https://julysun98.github.io/atypical_dataset。
原文摘要 · Abstract (English)
Humans usually show exceptional generalisation and discovery ability in the open world, when being shown uncommon new concepts. Whereas most existing studies in the literature focus on common typical data from closed sets, open-world novel discovery is under-explored in videos. In this paper, we are interested in asking: What if atypical unusual videos are exposed in the learning process? To this end, we collect a new video dataset consisting of various types of unusual atypical data (e.g., sci-fi, animation, etc.). To study how such atypical data may benefit open-world learning, we feed them into the model training process for representation learning. Focusing on three key tasks in open-world learning: out-of-distribution (OOD) detection, novel category discovery (NCD), and zero-shot action recognition (ZSAR), we found that even straightforward learning approaches with atypical data consistently improve performance across various settings. Furthermore, we found that increasing the categorical diversity of the atypical samples further boosts OOD detection performance. Additionally, in the NCD task, using a smaller yet more semantically diverse set of atypical samples leads to better performance compared to using a larger but more typical dataset. In the ZSAR setting, the semantic diversity of atypical videos helps the model generalise better to unseen action classes. These observations in our extensive experimental evaluations reveal the benefits of atypical videos for visual representation learning in the open world, together with the newly proposed dataset, encouraging further studies in this direction. The project page is at: https://julysun98.github.io/atypical_dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。