用多模态数据和伪缺失填补提升青蛙计数与分布预测精度
FrogDeepSDM: Improving Frog Counting and Occurrence Prediction Using Multimodal Data and Pseudo-Absence Imputation
- 融合影像与表格数据,通过伪缺失补全增强模型输入
- 计数任务MAE从189降至29,分类准确率达84.9%
- 适合生态监测、物种分布建模及数据稀疏场景
监测物种分布对保护工作至关重要,可评估环境影响并制定有效保育策略。传统调查方法(如公民科学)覆盖有限且不完整。物种分布建模(SDM)利用出现记录与环境变量,预测大范围物种分布。本研究基于「EY - 2022 生物多样性挑战」数据,采用深度学习与数据填补技术提升蛙类(Anura)SDM精度。实验表明,数据平衡显著改善模型表现,青蛙计数的均方绝对误差(MAE)从189降至29。特征选择识别出关键环境因子,优化输入同时保持预测能力。多模态集成模型融合土地覆盖、NDVI等环境变量,优于单一模型,在未见区域也具强泛化能力。图像与表格数据融合提升了计数与生境分类性能,达到84.9%准确率,AUC为0.90。结果表明,多模态学习与数据预处理(如平衡与填补)可在数据稀疏或不完整时显著提升生态预测建模效果,推动更精准、可扩展的生物多样性监测。
原文摘要 · Abstract (English)
Monitoring species distribution is vital for conservation efforts, enabling the assessment of environmental impacts and the development of effective preservation strategies. Traditional data collection methods, including citizen science, offer valuable insights but remain limited in coverage and completeness. Species Distribution Modelling (SDM) helps address these gaps by using occurrence data and environmental variables to predict species presence across large regions. In this study, we enhance SDM accuracy for frogs (Anura) by applying deep learning and data imputation techniques using data from the "EY - 2022 Biodiversity Challenge." Our experiments show that data balancing significantly improved model performance, reducing the Mean Absolute Error (MAE) from 189 to 29 in frog counting tasks. Feature selection identified key environmental factors influencing occurrence, optimizing inputs while maintaining predictive accuracy. The multimodal ensemble model, integrating land cover, NDVI, and other environmental inputs, outperformed individual models and showed robust generalization across unseen regions. The fusion of image and tabular data improved both frog counting and habitat classification, achieving 84.9% accuracy with an AUC of 0.90. This study highlights the potential of multimodal learning and data preprocessing techniques such as balancing and imputation to improve predictive ecological modeling when data are sparse or incomplete, contributing to more precise and scalable biodiversity monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。