arXiv:2511.07877cs.CV2025-11AAAI被引 3

用统一框架生成多种视觉任务表示,突破单任务模型限制。

Visual Bridge: Universal Visual Perception Representations Generating

  • 将多任务表示生成建模为统一流匹配问题,通过可扩展嵌入桥接不同任务。
  • 在分类、检测、分割等5项任务上零样本表现媲美专用模型。
  • 适合追求通用视觉建模的科研与工程人员,推动跨任务迁移研究。

扩散模型在文本到图像生成、深度估计、光流等孤立视觉任务中取得显著进展,但普遍受限于“单任务-单模型”范式,严重制约其在多任务场景下的泛化与扩展能力。受大语言模型跨域泛化能力启发,我们提出基于流匹配的通用视觉感知框架,可生成跨多任务的多样化视觉表示。该方法将图像块标记到任务特定表示的过程建模为统一的流匹配问题,而非独立生成或回归任务。通过强自监督基础模型作为锚点,并引入多尺度循环任务嵌入机制,学习一个通用速度场以弥合异构任务间的差距,支持高效灵活的表示迁移。在分类、检测、分割、深度估计和图文检索等任务上的大量实验表明,该模型在零样本和微调设置下均达到有竞争力的性能,优于先前通用模型及多个专用模型。消融实验进一步验证了框架的鲁棒性、可扩展性与泛化能力。本工作标志着迈向通用视觉感知的重要一步,为未来通用视觉建模研究提供坚实基础。

原文摘要 · Abstract (English)

Recent advances in diffusion models have achieved remarkable success in isolated computer vision tasks such as text-to-image generation, depth estimation, and optical flow. However, these models are often restricted by a ``single-task-single-model'' paradigm, severely limiting their generalizability and scalability in multi-task scenarios. Motivated by the cross-domain generalization ability of large language models, we propose a universal visual perception framework based on flow matching that can generate diverse visual representations across multiple tasks. Our approach formulates the process as a universal flow-matching problem from image patch tokens to task-specific representations rather than an independent generation or regression problem. By leveraging a strong self-supervised foundation model as the anchor and introducing a multi-scale, circular task embedding mechanism, our method learns a universal velocity field to bridge the gap between heterogeneous tasks, supporting efficient and flexible representation transfer. Extensive experiments on classification, detection, segmentation, depth estimation, and image-text retrieval demonstrate that our model achieves competitive performance in both zero-shot and fine-tuned settings, outperforming prior generalist and several specialist models. Ablation studies further validate the robustness, scalability, and generalization of our framework. Our work marks a significant step towards general-purpose visual perception, providing a solid foundation for future research in universal vision modeling.

通用视觉流匹配多任务表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。