arXiv:2605.06086cs.CV2026-05

用低秩超网络统一处理多模态图像缺失,效果超越现有方法。

LARGO: Low-Rank Hypernetwork for Handling Missing Modalities

  • 用张量分解压缩多种缺失模态的模型,仅用一个网络实现
  • 在脑肿瘤和中风数据集上47/52配置领先,平均分割提升2.53%
  • 方法可推广至非医疗数据,适合跨模态任务研究者

多模态图像分析中处理缺失模态是一个重要挑战,现有方法通常依赖复杂架构且难以迁移。本文提出LARGO,一种在权重空间操作的超网络,通过经典多线性(CP)张量分解,将$2^N-1$种针对不同缺失场景的专用模型压缩为单一网络。在BraTS 2018(4模态,15种缺失场景)和ISLES 2022(3模态,7种缺失场景)上的实验证明,该方法在52种配置中47次排名第一,平均Dice分数分别较最优基线(mmFormer、M$^{3}$AE、ShaSpec、SimMLM)提升+0.68%和+2.53%。在avMNIST上的初步实验表明,LARGO可能适用于非医疗领域的异构模态数据。

原文摘要 · Abstract (English)

Addressing missing modalities is an important challenge in multimodal image analysis and often relies on complex architectures that do not transfer easily to different datasets without architectural modifications or hyperparameter tuning. While most existing methods tackle this problem in feature space by engineering representations that are robust to missing inputs, we instead operate in weight space. We propose LARGO, a hypernetwork that compresses the $2^N-1$ dedicated missing-modality models into a single network by modelling the convolutional weights using the Canonical Polyadic (CP) tensor decomposition. Extensive experimental validation on BraTS 2018 (4 modalities, 15 scenarios) and ISLES 2022 (3 modalities, 7 scenarios) shows that our method ranks first in 47 out of 52 configurations, achieving average Dice improvements of +0.68$\%$ and +2.53$\%$ over state-of-the-art baselines (mmFormer, M$^{3}$AE, ShaSpec, SimMLM). A proof-of-concept experiment on avMNIST suggests that LARGO may extend beyond medical imaging to heterogeneous non-medical modalities.

多模态缺失模态低秩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。