arXiv:2503.03794cs.LGcs.AI2025-03中稿 · the 2025 IEEE Conf…被引 1

用合成数据提升藻华检测模型性能,避免真实数据不足问题。

Synthetic Data Augmentation for Enhancing Harmful Algal Bloom Detection with Machine Learning

  • 用高斯耦合生成包含水温、盐度等特征的合成数据
  • 中等规模合成数据使误差降低至0.1850(原为0.4706)
  • 适合需要低成本数据增强的环境监测研究者

有害藻华(HABs)对水生生态系统和公共健康构成严重威胁,全球造成巨大经济损失。早期检测至关重要,但常受限于高质量训练数据的匮乏。本研究探讨利用高斯耦合生成合成数据以增强基于机器学习的藻华检测系统。基于水温、盐度和紫外辐射等环境特征,生成了样本量为100至1000的合成数据集,目标变量为校正后的叶绿素a浓度。实验表明,适度的合成数据增强显著提升模型性能(均方根误差从0.4706降至0.1850,p < 0.001)。然而,过度使用合成数据会引入噪声,降低预测准确性,强调需平衡数据增强程度。研究结果表明,合成数据在提升藻华监测能力方面具有潜力,为早期预警与生态及公共卫生风险缓解提供可扩展、低成本的方法。

原文摘要 · Abstract (English)

Harmful Algal Blooms (HABs) pose severe threats to aquatic ecosystems and public health, resulting in substantial economic losses globally. Early detection is crucial but often hindered by the scarcity of high-quality datasets necessary for training reliable machine learning (ML) models. This study investigates the use of synthetic data augmentation using Gaussian Copulas to enhance ML-based HAB detection systems. Synthetic datasets of varying sizes (100-1,000 samples) were generated using relevant environmental features$\unicode{x2015}$water temperature, salinity, and UVB radiation$\unicode{x2015}$with corrected Chlorophyll-a concentration as the target variable. Experimental results demonstrate that moderate synthetic augmentation significantly improves model performance (RMSE reduced from 0.4706 to 0.1850; $p < 0.001$). However, excessive synthetic data introduces noise and reduces predictive accuracy, emphasizing the need for a balanced approach to data augmentation. These findings highlight the potential of synthetic data to enhance HAB monitoring systems, offering a scalable and cost-effective method for early detection and mitigation of ecological and public health risks.

藻华检测合成数据机器学习环境监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。