arXiv:2409.01633cs.LGcs.AI2024-09

通过特征增强与重构提升视觉分类性能

SleepNet and DreamNet: Enriching and Reconstructing Representations for Consolidated Visual Classification

  • 用预训练编码器结合监督学习增强特征表示
  • 通过编码解码框架重建隐藏状态,深化表征
  • 在多个数据集上超越现有方法,适合表征学习研究者

在视觉理解任务中,有效融合丰富特征表示与鲁棒分类机制仍是一大挑战。本文提出两种新型深度学习模型:SleepNet 和 DreamNet,旨在通过特征增强与重构策略提升表征利用效率。SleepNet 将预训练编码器获得的特征与监督学习结合,实现更强、更鲁棒的特征学习。在此基础上,DreamNet 引入预训练编码解码框架,对隐藏状态进行重建,从而实现视觉表征的深层整合与优化。实验表明,所提模型在多个基准上均显著优于现有最先进方法,验证了增强与重构策略的有效性。

原文摘要 · Abstract (English)

An effective integration of rich feature representations with robust classification mechanisms remains a key challenge in visual understanding tasks. This study introduces two novel deep learning models, SleepNet and DreamNet, which are designed to improve representation utilization through feature enrichment and reconstruction strategies. SleepNet integrates supervised learning with representations obtained from pre-trained encoders, leading to stronger and more robust feature learning. Building on this foundation, DreamNet incorporates pre-trained encoder decoder frameworks to reconstruct hidden states, allowing deeper consolidation and refinement of visual representations. Our experiments show that our models consistently achieve superior performance compared with existing state-of-the-art methods, demonstrating the effectiveness of the proposed enrichment and reconstruction approaches.

表征学习深度学习视觉分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。