研究发现:新类别发现效果随已知类别增多而提升,但存在饱和点。
How many classes do we need to see for novel class discovery?
- 用可控的dSprites数据集研究新类发现的影响因素。
- 已知类别超过一定数量后,发现性能不再提升,呈现收益递减。
- 为实际应用提供成本效益参考,适合关注模型泛化能力的研究者。
新类别发现对机器学习模型适应不断变化的真实世界数据至关重要,应用场景涵盖科学发现至机器人技术。然而,这些数据集包含复杂且纠缠的变量,使得系统性研究类别发现变得困难。因此,关于何时何因新类别发现更易成功等基本问题仍未解答。为此,我们提出一个基于程序生成变量的dSprites数据集的简单受控实验框架,以探究影响发现成功的因素。重点研究已知/未知类别数量与发现性能的关系,以及已知类别覆盖度对新类别发现的影响。实验结果表明,已知类别数量带来的益处存在饱和点,超过后发现性能趋于平稳。不同设置下的收益递减模式为从业者提供了成本效益分析依据,并为未来在复杂真实数据集上开展更严谨的类别发现研究提供了起点。
原文摘要 · Abstract (English)
Novel class discovery is essential for ML models to adapt to evolving real-world data, with applications ranging from scientific discovery to robotics. However, these datasets contain complex and entangled factors of variation, making a systematic study of class discovery difficult. As a result, many fundamental questions are yet to be answered on why and when new class discoveries are more likely to be successful. To address this, we propose a simple controlled experimental framework using the dSprites dataset with procedurally generated modifying factors. This allows us to investigate what influences successful class discovery. In particular, we study the relationship between the number of known/unknown classes and discovery performance, as well as the impact of known class 'coverage' on discovering new classes. Our empirical results indicate that the benefit of the number of known classes reaches a saturation point beyond which discovery performance plateaus. The pattern of diminishing return across different settings provides an insight for cost-benefit analysis for practitioners and a starting point for more rigorous future research of class discovery on complex real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。