arXiv:2510.20540cs.LG2025-10被引 4

提出新型去中心化多模态对齐框架,解决实际场景中模态不冗余的问题。

SheafAlign: A Sheaf-theoretic Framework for Decentralized Multimodal Alignment

  • 用层叠结构建模模态间两两关系,替代传统单一空间对齐
  • 在多模态传感数据上实现零样本泛化与缺模态鲁棒性
  • 通信成本比现有方法低50%,适合分布式系统部署

传统多模态对齐方法假设所有模态间存在相互冗余,这一假设在真实分布场景下失效。本文提出SheafAlign,一种基于层叠理论的去中心化多模态对齐框架,将单一空间对齐替换为多个对比空间。该方法通过层叠结构建模模态间的成对关系,并采用去中心化对比学习目标进行训练。SheafAlign克服了先前方法对全模态冗余的依赖,同时保留共享与独特信息。在多模态传感数据集上的实验表明,其在零样本泛化、跨模态对齐和缺失模态鲁棒性方面表现优异,通信成本比当前最优基线降低50%。

原文摘要 · Abstract (English)

Conventional multimodal alignment methods assume mutual redundancy across all modalities, an assumption that fails in real-world distributed scenarios. We propose SheafAlign, a sheaf-theoretic framework for decentralized multimodal alignment that replaces single-space alignment with multiple comparison spaces. This approach models pairwise modality relations through sheaf structures and leverages decentralized contrastive learning-based objectives for training. SheafAlign overcomes the limitations of prior methods by not requiring mutual redundancy among all modalities, preserving both shared and unique information. Experiments on multimodal sensing datasets show superior zero-shot generalization, cross-modal alignment, and robustness to missing modalities, with 50\% lower communication cost than state-of-the-art baselines.

多模态对齐去中心化层叠理论通信效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。