arXiv:2608.21281cs.CV2026-08

构建首个野外鱼类行为识别数据集,揭示模型在真实海洋环境中的巨大差距

WildFin: An In-the-Wild Dataset for Fish Behavioral Recognition

论文配图:WildFin: An In-the-Wild Dataset for Fish Behavioral Recognition
图 1 · 摘自论文原文
  • 由生态学家采集标注,覆盖固定摄像头与动态拍摄双场景
  • 耗时1350小时野外工作,600小时专家标注,产出超200万帧标签
  • 首次系统对比静态与时空模型,指出现有算法远未满足实际需求

近年来野外技术发展带来海量生态视频数据,但专家标注成本高昂成为主要瓶颈。尽管计算机视觉提供潜在解决方案,现有模型在复杂海洋环境中表现仍差。为分析此类失败,我们提出WildFin,一个由生态学家采集并标注的鱼类行为识别新基准。该数据集涵盖两类真实场景:固定摄像头监测鱼群与动态潜水员追踪个体。数据集历时1350小时野外作业、600小时专家标注,最终生成9小时行为数据,包含超过200万帧逐帧标签。我们对现代视觉基础模型进行基准测试,量化静态与时空架构间的权衡,揭示当前模型能力与真实水下行为分析需求之间仍存在显著差距。

原文摘要 · Abstract (English)

Recent advances in field technology have led to a massive influx of in-the-wild video data for ecological science. The primary bottleneck in leveraging this data is the high cost of expert annotation. While computer vision offers a potential solution, current models frequently fail when deployed in complex marine environments. To characterize these failures, we introduce WildFin, a novel benchmark for fish behavior recognition collected and annotated by ecologists. WildFin spans two critical real-world paradigms: stationary cameras monitoring groups of fish and dynamic divers following individual subjects. The dataset represents a massive curation effort, involving 1,350 hours of fieldwork and 600 hours of expert annotation to produce 9 hours of behavioral data with over 2 million frame-by-frame labels. We benchmark modern vision foundation models and quantify tradeoffs between static and spatiotemporal architectures, revealing the substantial gap that remains between current model capabilities and the demands of real-world underwater behavioral analysis. Project website: https://team-wildfin.github.io/.

行为识别野外数据水下视觉生态监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。