arXiv:2502.00474cs.CVcs.LG2025-02被引 1

用视觉模型自动分析溪流连通性,准确率提升至90%

A framework for river connectivity classification using temporal image processing and attention based neural networks

  • 结合时序图像处理与注意力机制,自动识别溪流连通状态
  • 在未见过的站点数据上准确率达90%,较基础模型提升15个百分点
  • 适合生态监测、气候变化研究者使用,尤其关注低成本监测

测量河流与溪流的水体连通性对水资源管理至关重要。气候变暖导致极端天气频发,影响河网连通性。传统流量监测设备成本高且仅适用于大河,而溪流摄像头可低成本、高频次采集图像。但每年需人工筛选数万张图像进行标注。为此,我们构建了一个自动化溪流图像分类框架,包含三部分:(1)图像预处理,含七种质量过滤、基于植被的亮度方差降低、缩放与底部中心裁剪;(2)采用扩散模型生成增强图像以平衡数据;(3)使用带时序增强的视觉变换器模型进行分类。基于康涅狄格州能源与环境部2018–2020年采集并标注的数据集,该框架将75%的基础准确率提升至90%,证明时序图像处理与注意力模型组合在未见站点图像分类中有效。

原文摘要 · Abstract (English)

Measuring the connectivity of water in rivers and streams is essential for effective water resource management. Increased extreme weather events associated with climate change can result in alterations to river and stream connectivity. While traditional stream flow gauges are costly to deploy and limited to large river bodies, trail camera methods are a low-cost and easily deployed alternative to collect hourly data. Image capturing, however requires stream ecologists to manually curate (select and label) tens of thousands of images per year. To improve this workflow, we developed an automated instream trail camera image classification system consisting of three parts: (1) image processing, (2) image augmentation and (3) machine learning. The image preprocessing consists of seven image quality filters, foliage-based luma variance reduction, resizing and bottom-center cropping. Images are balanced using variable amount of generative augmentation using diffusion models and then passed to a machine learning classification model in labeled form. By using the vision transformer architecture and temporal image enhancement in our framework, we are able to increase the 75% base accuracy to 90% for a new unseen site image. We make use of a dataset captured and labeled by staff from the Connecticut Department of Energy and Environmental Protection between 2018-2020. Our results indicate that a combination of temporal image processing and attention-based models are effective at classifying unseen river connectivity images.

图像分类生态监测视觉模型时序处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。