无标签数据仍含人类先验,应明确披露监督来源。
Position: Unlabeled IS NOT Equal to No Human Supervision in Visual Learning

- 区分不同数据筛选与训练目标中的隐含人类先验
- 呼吁用更清晰术语替代笼统的'无监督'标签
- 适合关注方法可复现性与公平比较的研究者
本文指出,视觉学习中无标签并不等于无人类监督。当前计算机视觉许多方法依赖大规模无标签数据预训练,被统称为‘无监督’,但不同的数据筛选方式与训练目标实际嵌入了显著不同的先验知识。单一‘无监督’术语已无法反映这些差异。这种模糊性阻碍了不同假设下研究的公平比较。尽管领域持续发展,2021年后顶级会议中以‘无监督’为题的论文数量却明显下降。我们支持预训练作为现代视觉模型的基础,但主张社区需提升概念清晰度:作者应明确披露数据选择与学习目标中的先验信息,并指明学习流程中各环节所依赖的假设。建立标准化披露规范有助于改善学术沟通、确保公平比较,并维护无监督学习的方法多样性。
原文摘要 · Abstract (English)
This position paper argues that the absence of labels does not imply the absence of human supervision in visual learning, and urges the research community to identify sources of supervision more explicitly. Many recent methods in computer vision build upon representations learned from large-scale unlabeled data, and are therefore grouped under the same umbrella term ``unsupervised.'' However, different data curation schemes and training objectives embed substantially different human priors on which models rely, and we argue that one ``unsupervised'' umbrella term is no longer capturing these distinctions. This ambiguity makes it harder to compare unsupervised learning research conducted under different assumptions, coinciding with a sharp decline in papers titled with ``unsupervised'' in flagship computer vision conferences since 2021, despite continued growth of the field. While we fully embrace pre-training as a strong foundation for modern computer vision, we advocate for a community-level effort toward greater conceptual clarity: authors are encouraged to disclose priors in data selection and learning objectives, and to specify which components of a learning pipeline depend on which assumptions. Standardized disclosure practices can improve academic communication, ensure fairer comparisons, and preserve methodological diversity in unsupervised learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。