通过自适应检索增强数据,提升环境模型在稀疏和异常情况下的预测能力
Learning to Retrieve for Environmental Knowledge Discovery: An Augmentation-Adaptive Self-Supervised Learning Framework
- 用多层级对比损失训练场景编码器,学习不同环境间的相似性
- 在数据稀缺或极端条件下,预测精度提升显著,鲁棒性更强
- 适合生态监测、气候变化研究等需要跨区域泛化的领域
环境知识发现依赖标注的任务特定数据,但数据采集成本高昂。现有机器学习方法在数据稀疏或非典型条件下泛化能力差。为此,我们提出一种增强自适应自监督学习框架(A²SL),通过检索相关观测样本,增强目标生态系统的建模能力。具体而言,引入多层级成对学习损失,训练场景编码器以捕捉不同场景间的差异相似性;这些相似性驱动检索机制,从不同地点或时间周期补充目标场景数据。此外,为更好应对多变场景,尤其在传统模型表现不佳的异常或极端条件下,设计了增强自适应机制,通过针对性数据增强提升模型性能。以淡水生态系统为例,我们在真实湖泊中评估了A²SL对水温与溶解氧动态的建模效果。实验结果表明,A²SL显著提升了预测准确性,并增强了在数据稀疏及非典型场景下的鲁棒性。尽管本研究聚焦淡水生态系统,该框架可广泛适用于多种科学领域。
原文摘要 · Abstract (English)
The discovery of environmental knowledge depends on labeled task-specific data, but is often constrained by the high cost of data collection. Existing machine learning approaches usually struggle to generalize in data-sparse or atypical conditions. To this end, we propose an Augmentation-Adaptive Self-Supervised Learning (A$^2$SL) framework, which retrieves relevant observational samples to enhance modeling of the target ecosystem. Specifically, we introduce a multi-level pairwise learning loss to train a scenario encoder that captures varying degrees of similarity among scenarios. These learned similarities drive a retrieval mechanism that supplements a target scenario with relevant data from different locations or time periods. Furthermore, to better handle variable scenarios, particularly under atypical or extreme conditions where traditional models struggle, we design an augmentation-adaptive mechanism that selectively enhances these scenarios through targeted data augmentation. Using freshwater ecosystems as a case study, we evaluate A$^2$SL in modeling water temperature and dissolved oxygen dynamics in real-world lakes. Experimental results show that A$^2$SL significantly improves predictive accuracy and enhances robustness in data-scarce and atypical scenarios. Although this study focuses on freshwater ecosystems, the A$^2$SL framework offers a broadly applicable solution in various scientific domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。