arXiv:2605.03348cs.LGcs.AI2026-05

通过专家选择与稀疏化,构建可选的语义模块化多模态表示。

Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts

论文配图:Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts
图 1 · 摘自论文原文
  • 将多模态输入分解为概念级专家,按任务动态路由
  • 在四个基准上提升准确率,中等稀疏度时性能最佳
  • 适合需要高效、可解释多模态模型的研究者

我们提出S3(Specialization, Selection, Sparsification)框架,从结构视角重新思考多模态学习。不同于将所有信号编码为固定嵌入,S3将多模态输入分解为共享隐空间中的语义专家,并根据任务需求选择性路由。专业化在共享隐空间中形成概念级专家,选择性适配任务特定路由,稀疏化则剪除低效路径,生成紧凑、信息量最小的表示。在四个MultiBench基准上,S3提升准确率并表现出一致的倒U型稀疏性-性能关系,峰值性能出现在中等稀疏度。结果表明,将多模态表示结构化为可选语义组件,是对比学习或InfoMax驱动方法的一种实用且有原则的替代方案。

原文摘要 · Abstract (English)

We propose S3 (Specialization, Selection, Sparsification), a framework that rethinks multimodal learning through a structural perspective. Instead of encoding all signals into a fixed embedding, S3 decomposes multimodal inputs into semantic experts and selectively routes them for each task. Specialization forms concept-level experts in a shared latent space, Selection adapts routing for task-specific needs, and Sparsification prunes low-utility paths to yield compact, information-minimal representations. Across four MultiBench benchmarks, S3 improves accuracy and shows a consistent reverse U-shaped sparsity-performance trend, with peak performance at intermediate sparsity. These results suggest that structuring multimodal representations as selectable semantic components provides a practical and principled alternative to contrastive learning or InfoMax-driven approaches.

多模态专家混合稀疏表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。