arXiv:2509.25278cs.LGstat.ML2025-09NeurIPS被引 12

MAESTRO动态融合多模态时序数据,应对传感器缺失与模态不均衡挑战。

MAESTRO : Adaptive Sparse Attention and Robust Learning for Multimodal Dynamic Time Series

  • 基于任务相关性动态建模模态间交互,自适应分配注意力预算。
  • 在完整数据下相比最优方法提升4%~8%,缺失40%模态时仍领先9%。
  • 适合医疗监测等真实场景中存在传感器故障的多模态时序分析。

从临床医疗到日常生活,多模态连续传感器监测在智能决策中展现出巨大潜力,但也面临诸多挑战。本文提出MAESTRO框架,克服现有方法三大局限:(1) 依赖单一主模态对齐,(2) 采用成对模态建模,(3) 假设所有模态观测完整。这些限制阻碍了其在真实多模态时序场景中的应用,因主模态先验常不明确、模态数量大(成对建模不可行)、传感器故障导致任意缺失。MAESTRO通过任务相关性驱动的动态模内与模间交互,结合符号化分词与自适应注意力预算,构建长序列多模态表示,并采用稀疏跨模态注意力处理。生成的跨模态特征经稀疏Mixture-of-Experts(MoE)机制路由,实现不同模态组合下的黑箱专业化。在四个覆盖三个应用领域的数据集上,对比10个基线,完整观测下平均相对提升4%和8%(优于最佳多模态与多变量方法);部分观测下(最多40%模态缺失),平均提升9%。进一步分析表明,其稀疏且模态感知的设计具有鲁棒性与高效性。

原文摘要 · Abstract (English)

From clinical healthcare to daily living, continuous sensor monitoring across multiple modalities has shown great promise for real-world intelligent decision-making but also faces various challenges. In this work, we introduce MAESTRO, a novel framework that overcomes key limitations of existing multimodal learning approaches: (1) reliance on a single primary modality for alignment, (2) pairwise modeling of modalities, and (3) assumption of complete modality observations. These limitations hinder the applicability of these approaches in real-world multimodal time-series settings, where primary modality priors are often unclear, the number of modalities can be large (making pairwise modeling impractical), and sensor failures often result in arbitrary missing observations. At its core, MAESTRO facilitates dynamic intra- and cross-modal interactions based on task relevance, and leverages symbolic tokenization and adaptive attention budgeting to construct long multimodal sequences, which are processed via sparse cross-modal attention. The resulting cross-modal tokens are routed through a sparse Mixture-of-Experts (MoE) mechanism, enabling black-box specialization under varying modality combinations. We evaluate MAESTRO against 10 baselines on four diverse datasets spanning three applications, and observe average relative improvements of 4% and 8% over the best existing multimodal and multivariate approaches, respectively, under complete observations. Under partial observations -- with up to 40% of missing modalities -- MAESTRO achieves an average 9% improvement. Further analysis also demonstrates the robustness and efficiency of MAESTRO's sparse, modality-aware design for learning from dynamic time series.

多模态时序建模稀疏注意力传感器缺失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。