arXiv:2509.25678cs.LG2025-09被引 1

用时间依赖关系指导专家路由,让多模态模型更好捕捉传感器间延迟交互。

Massively Multimodal Foundation Models: A Framework for Capturing Interactions with Specialized Mixture-of-Experts

  • 根据模态间时间延迟设计交互感知路由
  • 在医疗、行为识别等任务上显著提升性能
  • 适合处理大量异构传感器数据的场景

现代应用越来越多地涉及多种异构输入流,如临床传感器、可穿戴设备数据、影像和文本,每种数据具有不同的测量模型、采样率和噪声特性。我们定义此为大规模多模态场景,每个传感器构成一个独立模态。随着模态数量增加,捕捉其复杂的时变交互(如传感器间的延迟生理级联)变得至关重要但极具挑战。混合专家(MoE)架构天然适用于此场景,因其稀疏路由机制可高效扩展至多个模态。然而,现有MoE架构仅基于相似性路由,忽视了模态间的丰富时序依赖,导致无法捕捉延迟的跨模态效应,进而造成专家分工不优与性能下降。我们提出一种框架,显式量化模态对在多个离散时间区间内的时序依赖关系(即事件在某一输入流中发生与其在另一流中显现影响之间的延迟),并利用这些信息引导MoE路由。交互感知路由器根据交互类型将令牌分发至专用专家,使专家能够学习通用的交互处理能力。在医疗、活动识别和情感计算基准上的实验表明,该方法带来显著性能提升,并生成与领域知识一致的可解释路由模式。

原文摘要 · Abstract (English)

Modern applications increasingly involve many heterogeneous input streams, such as clinical sensors, wearable device data, imaging, and text, each with distinct measurement models, sampling rates, and noise characteristics. We define this as massively multimodal setting, where each sensor constitutes a separate modality. As modality counts grow, capturing their complex, time-varying interactions such as delayed physiological cascades between sensors, has becomes essential yet challenging. Mixture-of-Experts (MoE) architectures are naturally suited for this setting since their sparse routing mechanism enables efficient scaling across many modalities. However, existing MoE architectures route tokens based on similarity alone, overlooking the rich temporal dependencies across modalities: this prevents the model from capturing delayed cross-modal effects, leading to suboptimal expert specialization and reduced accuracy. We propose a framework that explicitly quantifies temporal dependencies between modality pairs across multiple discrete time intervals, defined as delays between an event in one input stream and its manifested effect in another, and uses these to guide MoE routing. A interaction-aware router dispatches tokens to specialized experts based on interaction type. This principled routing enables experts to learn generalizable interaction-processing skills. Experiments across healthcare, activity recognition, and affective computing benchmarks demonstrate substantial performance gains and interpretable routing patterns aligned with domain knowledge.

多模态MoE时序建模医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。