arXiv:2602.04021cs.LGq-bio.QM2026-02

弱配对多模态数据中,通过分组对比学习提升表示能力。

Group Contrastive Learning for Weakly Paired Multimodal Data

  • 设计分组对比损失,利用共享扰动标签建模弱配对关系。
  • 在真实单细胞数据上表现优于或持平现有方法,支持跨模态匹配与补全。
  • 提供系统评估框架,可检验不同对齐器在多种扰动场景下的鲁棒性。

我们提出 GROOVE,一种针对高内容扰动数据的半监督多模态表征学习方法,其中不同模态样本仅通过共享扰动标签弱配对,缺乏直接对应关系。核心贡献是 GroupCLIP——一种新型分组级对比损失,弥合了适用于成对跨模态数据的 CLIP 与适用于单模态监督对比学习的 SupCon 之间的基本差距。我们将 GroupCLIP 与实时反向翻译自编码器框架结合,促进跨模态纠缠表示,同时保持共享潜在空间内的分组一致性。关键的是,我们引入了一个综合组合评估框架,系统评估多个最优传输对齐器的表现,解决了现有评估策略的关键局限。该框架包含新设计的模拟实验,系统变化共享与模态特异性扰动效应,实现方法鲁棒性的合理评估。组合基准测试表明,尚无对齐器在所有设置或模态对中全面占优。在模拟和两个真实单细胞基因扰动数据集上,GROOVE 在下游跨模态匹配与补全任务中表现与或优于现有方法。消融实验表明,GroupCLIP 是性能提升的关键组件。结果凸显了在仅存在弱配对的情况下,利用分组约束进行有效多模态表征学习的重要性。

原文摘要 · Abstract (English)

We present GROOVE, a semi-supervised multi-modal representation learning approach for high-content perturbation data where samples across modalities are weakly paired through shared perturbation labels but lack direct correspondence. Our primary contribution is GroupCLIP, a novel group-level contrastive loss that bridges the gap between CLIP for paired cross-modal data and SupCon for uni-modal supervised contrastive learning, addressing a fundamental gap in contrastive learning for weakly-paired settings. We integrate GroupCLIP with an on-the-fly backtranslating autoencoder framework to encourage cross-modally entangled representations while maintaining group-level coherence within a shared latent space. Critically, we introduce a comprehensive combinatorial evaluation framework that systematically assesses representation learners across multiple optimal transport aligners, addressing key limitations in existing evaluation strategies. This framework includes novel simulations that systematically vary shared versus modality-specific perturbation effects enabling principled assessment of method robustness. Our combinatorial benchmarking reveals that there is not yet an aligner that uniformly dominates across settings or modality pairs. Across simulations and two real single-cell genetic perturbation datasets, GROOVE performs on par with or outperforms existing approaches for downstream cross-modal matching and imputation tasks. Our ablation studies demonstrate that GroupCLIP is the key component driving performance gains. These results highlight the importance of leveraging group-level constraints for effective multi-modal representation learning in scenarios where only weak pairing is available.

多模态学习对比学习弱配对单细胞数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。