用多任务学习提升汽车品牌与型号的层级分类效果
An Analysis of Multi-Task Architectures for the Hierarchic Multi-Label Problem of Vehicle Model and Make Classification
- 设计并对比并行与级联式多任务架构
- 在StanfordCars和CompCars上均提升准确率,尤其在CompCars上显著
- 适合研究层级分类、多任务学习的工程师与研究人员
世界中的信息大多具有层次结构,但多数深度学习方法未利用这一语义丰富的结构。研究表明人类学习能受益于信息的层级组织,智能模型亦可通过多任务学习实现类似优势。本文针对汽车品牌与型号的层级多标签分类问题,分析多任务学习的优劣。比较了并行与级联两种多任务架构,在CNN与Transformer等模型上,调整丢弃率与损失权重,评估其在斯坦福车辆数据集(StanfordCars)与CompCars上的表现。实验表明,多任务范式在两个数据集上均有效,几乎所有情况下均提升所研究CNN的性能;尤其在CompCars上,两类模型均获得显著改进。
原文摘要 · Abstract (English)
Most information in our world is organized hierarchically; however, many Deep Learning approaches do not leverage this semantically rich structure. Research suggests that human learning benefits from exploiting the hierarchical structure of information, and intelligent models could similarly take advantage of this through multi-task learning. In this work, we analyze the advantages and limitations of multi-task learning in a hierarchical multi-label classification problem: car make and model classification. Considering both parallel and cascaded multi-task architectures, we evaluate their impact on different Deep Learning classifiers (CNNs, Transformers) while varying key factors such as dropout rate and loss weighting to gain deeper insight into the effectiveness of this approach. The tests are conducted on two established benchmarks: StanfordCars and CompCars. We observe the effectiveness of the multi-task paradigm on both datasets, improving the performance of the investigated CNN in almost all scenarios. Furthermore, the approach yields significant improvements on the CompCars dataset for both types of models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。