用维基文本和众包观测数据,零样本估算物种分布范围。
Combining Observational Data and Language for Species Range Estimation
- 融合维基文本与众包观测数据,构建跨模态统一表征空间。
- 零样本预测性能媲美需数十次观测的传统方法。
- 适合数据稀缺物种,可作先验提升小样本建模效果。
物种分布图(SRMs)是生态学、保护生物学与环境管理的重要工具。传统方法依赖环境变量和高质量物种观测数据,但常因地理障碍与资源限制难以获取。本文提出新方法:结合数百万条公民科学观测数据与维基百科中关于栖息地偏好和分布描述的文本,覆盖数万种物种。通过将位置、物种与文本映射至统一嵌入空间,实现全球尺度上丰富的空间协变量学习,并支持仅凭文本进行零样本分布估计。在保留物种上的评估显示,零样本模型显著优于基线,性能接近使用数十条观测数据的传统方法。该方法在结合真实观测数据时亦能作为强先验,以更少数据实现更高精度的分布预测。我们通过大量定量与定性分析验证了所学表示在分布估计及其他空间任务中的有效性。
原文摘要 · Abstract (English)
Species range maps (SRMs) are essential tools for research and policy-making in ecology, conservation, and environmental management. However, traditional SRMs rely on the availability of environmental covariates and high-quality species location observation data, both of which can be challenging to obtain due to geographic inaccessibility and resource constraints. We propose a novel approach combining millions of citizen science species observations with textual descriptions from Wikipedia, covering habitat preferences and range descriptions for tens of thousands of species. Our framework maps locations, species, and text descriptions into a common space, facilitating the learning of rich spatial covariates at a global scale and enabling zero-shot range estimation from textual descriptions. Evaluated on held-out species, our zero-shot SRMs significantly outperform baselines and match the performance of SRMs obtained using tens of observations. Our approach also acts as a strong prior when combined with observational data, resulting in more accurate range estimation with less data. We present extensive quantitative and qualitative analyses of the learned representations in the context of range estimation and other spatial tasks, demonstrating the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。