arXiv:2606.00275cs.CVcs.AI2026-06

提出不对称专家架构,提升视觉语言模型的推理准确性和效率。

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models

论文配图:Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models
图 1 · 摘自论文原文
  • 分三类专家:模态内、双曲跨模态、证据优先语言专家
  • 在幻觉敏感任务上最高提升3.8%,平均比MoE高1.5%
  • 激活参数减少25.45%,适合资源受限场景

大型视觉语言模型(LVLM)通过大规模架构和训练展现出卓越性能。近期研究将混合专家(MoE)引入LVLM以提升计算效率。然而,现有方法对视觉与语言模态采用对称结构,忽略了二者处理上的固有不对称性。这种不对称导致两大问题:第一,文本与视觉呈层次而非并行关系,文本查询通常仅描述完整视觉场景的部分内容,欧几里得专家空间难以建模此类包含结构;第二,深层语言专家逐步从基于证据的处理转向参数记忆依赖,失去对输入视觉与语言信息的锚定。为此,我们提出AsyMoE,通过三类专用专家显式建模此不对称性:模态内专家处理模态特异性任务;双曲跨模态专家利用负曲率几何捕捉层次化跨模态关系;证据优先语言专家抑制参数记忆激活,保持全程上下文接地。大量实验表明,AsyMoE相较基线方法持续提升性能,平均优于MoE变体1.5%,在幻觉敏感任务上最高提升3.8%。相较于密集模型,AsyMoE激活参数减少25.45%。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have demonstrated impressive performance on multimodal tasks through scaled architectures and extensive training. Recent studies introduce Mixture of Experts (MoE) into LVLMs for improved computational efficiency. However, existing MoE approaches treat visual and linguistic modalities with symmetric architectures, overlooking the inherent asymmetry in how these two modalities are processed. This asymmetry causes two critical issues. First, text and vision form hierarchical rather than parallel relationships, as text queries typically describe partial aspects of complete visual scenes. Euclidean expert space struggles to encode such containment structures. Second, language experts in deeper layers progressively shift from evidence-based processing to parametric memory dependence, losing grounding in the provided visual and linguistic information. To address these issues, we propose AsyMoE, a novel architecture that explicitly models this asymmetry through three specialized expert groups. Intra-modality experts handle modality-specific processing. Hyperbolic inter-modality experts capture hierarchical cross-modal relationships through negative curvature geometry. Evidence-priority language experts suppress parametric memory activation and maintain contextual grounding throughout network depth. Extensive experiments demonstrate that AsyMoE achieves consistent improvements over baseline methods, with average gains of 1.5\% over MoE variants and up to 3.8\% on hallucination-sensitive tasks. AsyMoE activates 25.45\% fewer parameters compared to dense models.

视觉语言模型混合专家双曲几何推理效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。