arXiv:2412.10730cs.CV2024-12

通过聚类掩码与多任务预训练,提升xLSTM在视觉任务中的表现。

MAL: Cluster-Masked and Multi-Task Pretraining for Enhanced xLSTM Vision Performance

  • 设计聚类掩码机制,优化局部特征捕捉与图像扫描效率。
  • 融合自回归、深度估计与分割任务,实现统一预训练。
  • 在多种视觉任务中超越监督模型,释放xLSTM的可扩展潜力。

传统长短期记忆(LSTM)网络在视觉任务中面临难以扩展和捕捉复杂依赖的问题。xLSTM通过引入指数门控和并行矩阵记忆结构,提升了性能与可扩展性。然而,xLSTM在视觉计算中的潜力尚未充分挖掘,尤其是在利用自回归技术进行特征提取方面。本文提出MAL(Cluster-Masked and Multi-Task Pretraining for Enhanced xLSTM Vision Performance),一种新型框架,通过创新的预训练策略增强xLSTM能力。我们提出一种聚类掩码方法,显著提升局部特征捕捉能力并优化图像扫描效率。此外,通用编码器-解码器预训练方法整合了图像自回归、深度估计与图像分割等多任务,增强模型在多样视觉任务中的适应性与鲁棒性。实验表明,MAL超越传统监督模型,充分释放xLSTM的可扩展性,树立了视觉任务性能新基准。

原文摘要 · Abstract (English)

The Long Short-Term Memory (LSTM) networks have traditionally faced challenges in scaling and effectively capturing complex dependencies in visual tasks. The xLSTM architecture has emerged to address these limitations, incorporating exponential gating and a parallel matrix memory structure to enhance performance and scalability. Despite these advancements, the potential of xLSTM in visual computing has not been fully realized, particularly in leveraging autoregressive techniques for improved feature extraction. In this paper, we introduce MAL (Cluster-Masked and Multi-Task Pretraining for Enhanced xLSTM Vision Performance), a novel framework that enhances xLSTM's capabilities through innovative pretraining strategies. We propose a cluster-masked masking method that significantly improves local feature capture and optimizes image scanning efficiency. Additionally, our universal encoder-decoder pretraining approach integrates multiple tasks, including image autoregression, depth estimation, and image segmentation, thereby enhancing the model's adaptability and robustness across diverse visual tasks. Our experimental results demonstrate that MAL surpasses traditional supervised models and fully leverages the scaling potential of xLSTM, setting a new benchmark in visual task performance.

xLSTM视觉预训练多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。