arXiv:2602.08282cs.CVcs.AI2026-02被引 1

融合有标签与仅有存在数据,提升植物分布预测准确性。

Tighnari v2: Mitigating Label Noise and Distribution Shift in Multimodal Plant Distribution Prediction via Mixture of Experts and Weakly Supervised Learning

  • 用卫星影像地理覆盖生成伪标签,对齐遥感特征空间。
  • 基于专家混合机制,分区处理不同地理分布区域样本。
  • 在小样本和分布偏移场景下表现更优,适合生态监测应用。

大规模跨物种植物分布预测对生物多样性保护至关重要,但观测数据稀疏且存在偏差。存在-缺失(PA)数据标签准确无噪声,但获取成本高、数量有限;存在-仅(PO)数据覆盖广、时空信息丰富,但负样本噪声严重。本文提出一种多模态融合框架,充分利用两类数据优势。引入基于卫星影像地理覆盖的伪标签聚合策略,实现标签空间与遥感特征空间的地理对齐。模型架构上,采用Swin Transformer Base作为遥感图像主干,使用TabM网络提取表格特征,保留Temporal Swin Transformer处理时序数据,并设计可堆叠的串行三模态交叉注意力机制优化异构模态融合。实证分析发现PA训练与测试样本间存在显著地理分布偏移,直接混合训练易受PO数据噪声影响导致性能下降。为此,借鉴专家混合范式:按测试样本与PA样本的空间邻近性划分区域,针对不同区域使用分别训练的模型进行推理与后处理。在GeoLifeCLEF 2025数据集上的实验表明,该方法在PA数据覆盖率低且分布偏移明显的情况下仍能取得优越预测性能。

原文摘要 · Abstract (English)

Large-scale, cross-species plant distribution prediction plays a crucial role in biodiversity conservation, yet modeling efforts in this area still face significant challenges due to the sparsity and bias of observational data. Presence-Absence (PA) data provide accurate and noise-free labels, but are costly to obtain and limited in quantity; Presence-Only (PO) data, by contrast, offer broad spatial coverage and rich spatiotemporal distribution, but suffer from severe label noise in negative samples. To address these real-world constraints, this paper proposes a multimodal fusion framework that fully leverages the strengths of both PA and PO data. We introduce an innovative pseudo-label aggregation strategy for PO data based on the geographic coverage of satellite imagery, enabling geographic alignment between the label space and remote sensing feature space. In terms of model architecture, we adopt Swin Transformer Base as the backbone for satellite imagery, utilize the TabM network for tabular feature extraction, retain the Temporal Swin Transformer for time-series modeling, and employ a stackable serial tri-modal cross-attention mechanism to optimize the fusion of heterogeneous modalities. Furthermore, empirical analysis reveals significant geographic distribution shifts between PA training and test samples, and models trained by directly mixing PO and PA data tend to experience performance degradation due to label noise in PO data. To address this, we draw on the mixture-of-experts paradigm: test samples are partitioned according to their spatial proximity to PA samples, and different models trained on distinct datasets are used for inference and post-processing within each partition. Experiments on the GeoLifeCLEF 2025 dataset demonstrate that our approach achieves superior predictive performance in scenarios with limited PA coverage and pronounced distribution shifts.

植物分布多模态专家混合弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。