用森林引导的语义传输提升多模态数据对齐精度
Forest-Guided Semantic Transport for Label-Supervised Manifold Alignment
- 基于标签感知的森林结构构建语义关系,替代传统欧氏几何
- 在合成数据和单细胞数据上显著提升对应点恢复与标签迁移效果
- 适合生物信息学中的批次校正与生物学保守性分析场景
标签监督的流形对齐通过共享标签信息连接无监督与基于对应关系的范式。然而,现有方法多依赖欧氏几何建模域内关系,在特征与任务关联较弱时会引入噪声、产生语义误导,降低对齐质量。为此,我们提出FoSTA(Forest-guided Semantic Transport Alignment)——一种可扩展的对齐框架,利用森林诱导的几何结构去噪域内结构,恢复任务相关的流形后再进行对齐。FoSTA直接从标签引导的森林亲密度构建语义表示,并通过快速分层语义传输实现对齐,捕捉有意义的跨域关系。与多个基准方法的广泛对比表明,FoSTA在合成基准上提升了对应关系恢复与标签迁移性能,并在真实单细胞应用中表现出色,包括批次校正与生物学保守性分析。
原文摘要 · Abstract (English)
Label-supervised manifold alignment bridges the gap between unsupervised and correspondence-based paradigms by leveraging shared label information to align multimodal datasets. Still, most existing methods rely on Euclidean geometry to model intra-domain relationships. This approach can fail when features are only weakly related to the task of interest, leading to noisy, semantically misleading structure and degraded alignment quality. To address this limitation, we introduce FoSTA (Forest-guided Semantic Transport Alignment), a scalable alignment framework that leverages forest-induced geometry to denoise intra-domain structure and recover task-relevant manifolds prior to alignment. FoSTA builds semantic representations directly from label-informed forest affinities and aligns them via fast, hierarchical semantic transport, capturing meaningful cross-domain relationships. Extensive comparisons with established baselines demonstrate that FoSTA improves correspondence recovery and label transfer on synthetic benchmarks and delivers strong performance in practical single-cell applications, including batch correction and biological conservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。