arXiv:2602.14039cs.CLcs.LG2026-02

提出新聚合方法,让专家模型输出更准确保留几何结构。

Geometry-Preserving Aggregation for Mixture-of-Experts Embedding Models

  • 分离径向与角度分量,保持专家输出的球面结构
  • 在MTEB任务上提升性能,训练成本不变且稳定
  • 适合需要精确向量对齐的嵌入模型研究者

混合专家(MoE)嵌入模型通过加权线性求和融合专家输出,隐含假设嵌入空间具有线性子空间结构。但现代MoE嵌入模型的几何分析显示,专家输出位于共享超球面流形上,表现为极紧凑的范数和显著的角度分离。在此几何下,线性聚合会导致向内坍缩,扭曲向量模长与方向,降低嵌入可比性。为此,本文提出球面重心聚合(SBA),一种保持几何一致性的聚合算子,分离径向与角度成分,在不改变现有路由机制的前提下维持超球面结构。在大规模文本嵌入基准(MTEB)的语义相似度、聚类和重复问题检测等任务中,实验表明性能持续提升,训练成本相同且完全稳定。额外几何分析证实SBA有效防止聚合导致的坍缩,保持超球面一致性,凸显几何感知聚合在MoE嵌入架构中的重要性。

原文摘要 · Abstract (English)

Mixture-of-Experts (MoE) embedding models combine expert outputs using weighted linear summation, implicitly assuming a linear subspace structure in the embedding space. This assumption is shown to be inconsistent with the geometry of expert representations. Geometric analysis of a modern MoE embedding model reveals that expert outputs lie on a shared hyperspherical manifold characterized by tightly concentrated norms and substantial angular separation. Under this geometry, linear aggregation induces inward collapse toward the manifold interior, distorting vector magnitude and direction and reducing embedding comparability. To address this inconsistency, Spherical Barycentric Aggregation (SBA) is introduced as a geometry-preserving aggregation operator that separates radial and angular components to maintain hyperspherical structure while remaining fully compatible with existing routing mechanisms. Experiments on selected tasks from the Massive Text Embedding Benchmark (MTEB), including semantic similarity, clustering, and duplicate question detection, demonstrate consistent performance improvements with identical training cost and full stability. Additional geometric analyses confirm that SBA prevents aggregation-induced collapse and preserves hyperspherical consistency, highlighting the importance of geometry-aware aggregation in MoE embedding architectures.

MoE嵌入几何聚合向量对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。