arXiv:2603.03564cs.CV2026-03

让视觉模型跨模态协同推理,提升图像视频3D理解能力

Modeling Cross-vision Synergy for Unified Large Vision Model

  • 用动态路由专家系统实现多模态专精与双向交互
  • 在10个基准上平均性能超越基线超10%
  • 适合需要跨视觉模态融合的AI研究者

近期大视觉模型(LVMs)从单一模态设计转向统一架构,联合处理图像、视频和3D数据。然而现有统一LVM主要追求功能整合,忽视了跨视觉协同的核心目标:利用不同视觉模态间的互补先验进行推理。为此,我们提出PolyV,一种在架构与训练层面均实现跨视觉协同的统一LVM。架构上,PolyV采用稀疏专家混合(Mixture-of-Experts)结构,由动态模态路由器协调,使每个专家专注特定模态先验,同时支持模态间双向交互与相互优化。训练上,采用协同感知范式,结合模态特异性预训练与基于知识蒸馏及对象/关系级对齐的粗到细协同调优。在涵盖图像、视频和3D理解的10个基准测试中,包括需空间或时间先验的协同聚焦数据集,PolyV持续优于现有模型,平均性能超越其基线超过10%。整体而言,PolyV建立了一个具同步感知能力的统一视觉推理框架,推动真正协同的大视觉模型发展。

原文摘要 · Abstract (English)

Recent advances in large vision models (LVMs) have shifted from modality-specific designs toward unified architectures that jointly process images, videos, and 3D data. However, existing unified LVMs primarily pursue functional integration, while overlooking the deeper goal of cross-vision synergy: the ability to reason over complementary priors across visual modalities. To address this, we present PolyV, a unified LVM that achieves cross-vision synergy at both the architectural and training levels. Architecturally, PolyV adopts a sparse Mixture-of-Experts LVM coordinated by a dynamic modality router, allowing each expert to specialize in modality-specific priors while enabling bidirectional interaction and mutual refinement across modalities. Training-wise, a synergy-aware paradigm combines modality-specific pretraining with coarse-to-fine synergy tuning via knowledge distillation and object-/relation-level alignment. Extensive experiments on 10 benchmarks spanning image, video, and 3D understanding, including synergy-focused datasets requiring spatial or temporal priors, demonstrate that PolyV consistently outperforms existing models, achieving over 10% average improvement over its backbone. Overall, PolyV establishes a unified framework for synesthetic visual reasoning, advancing toward truly synergistic LVMs. Project page: https://sqwu.top/PolyV.

视觉模型多模态协同推理统一架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。