arXiv:2509.20481cs.CVcs.AI2025-09

构建统一神经空间,让多种视觉任务共享特征,提升效率与泛化能力。

Shared Neural Space: Unified Precomputed Feature Encoding for Multi-Task and Cross Domain Vision

  • 设计轻量级CNN架构的编码器-解码器,预计算跨任务通用特征。
  • 在去马赛克、去噪、深度估计和语义分割上实现高效多任务处理。
  • 适合部署在资源受限设备,支持跨领域视觉任务快速集成。

当前多数成像与视觉AI模型针对特定高精度任务定制,但在一系列模块化任务中效率低下,因每个任务需映射到不同的潜在空间。为解决此问题,我们提出一种通用神经空间(NS),通过编码器-解码器框架在视觉与成像任务间预计算特征。编码器学习具有变换感知能力的通用表示,使多个下游AI模块共享同一特征空间。该架构减少冗余,提升跨域漂移下的泛化能力,并为高效多任务视觉流水线奠定基础。此外,相比大型Transformer骨干网络,我们的骨干为轻量级CNN,适配更广泛的硬件平台。进一步实验表明,去马赛克、去噪、深度估计和语义分割等成像与视觉模块可在该神经空间中高效执行。

原文摘要 · Abstract (English)

The majority of AI models in imaging and vision are customized to perform on specific high-precision task. However, this strategy is inefficient for applications with a series of modular tasks, since each requires a mapping into a disparate latent domain. To address this inefficiency, we proposed a universal Neural Space (NS), where an encoder-decoder framework pre-computes features across vision and imaging tasks. Our encoder learns transformation aware, generalizable representations, which enable multiple downstream AI modules to share the same feature space. This architecture reduces redundancy, improves generalization across domain shift, and establishes a foundation for effecient multi-task vision pipelines. Furthermore, as opposed to larger transformer backbones, our backbone is lightweight and CNN-based, allowing for wider across hardware. We furthur demonstrate that imaging and vision modules, such as demosaicing, denoising, depth estimation and semantic segmentation can be performed efficiently in the NS.

多任务学习神经空间轻量级模型视觉任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。