用AI从相机陷阱图像中自动提取物种与环境信息,生成丰富生态报告。
Towards Context-Rich Automated Biodiversity Assessments: Deriving AI-Powered Insights from Camera Trap Data
- 分两阶段:先用YOLOv10-X定位分类物种,再用Phi-3.5视觉模型补充上下文。
- 可识别物种、植被类型、时间等信息,支持复杂查询与结构化报告生成。
- 适合生态保护者快速获取物种分布、行为及栖息地偏好等关键洞察。
相机陷阱为生态研究带来巨大机遇,但现有自动化图像分析方法常缺乏支持有效保护决策所需的上下文信息。本文提出一种集成方法,结合基于深度学习的视觉与语言模型,提升相机陷阱数据的生态报告能力。采用两阶段系统:首先使用YOLOv10-X在图像中定位并分类哺乳动物和鸟类;随后利用Phi-3.5-vision-instruct模型读取边界框标签,识别难以分类的物体,弥补其局限性。此外,Phi-3.5还能检测植被类型、时间等更广泛的变量,为YOLO的物种检测结果提供丰富的生态与环境上下文。整合输出后,通过模型的自然语言系统回答复杂问题,并采用检索增强生成(RAG)技术,引入外部信息如物种体重、濒危等级(IUCN状态),这些无法直接从视觉分析获得。最终自动生成结构化报告,为生物多样性利益相关方提供关于物种丰度、分布、动物行为和栖息地选择的深入洞察。该方法生成具有上下文丰富性的叙述,不仅减少人工工作量,还支持保护决策的及时性,有助于推动管理从被动响应转向主动干预。
原文摘要 · Abstract (English)
Camera traps offer enormous new opportunities in ecological studies, but current automated image analysis methods often lack the contextual richness needed to support impactful conservation outcomes. Here we present an integrated approach that combines deep learning-based vision and language models to improve ecological reporting using data from camera traps. We introduce a two-stage system: YOLOv10-X to localise and classify species (mammals and birds) within images, and a Phi-3.5-vision-instruct model to read YOLOv10-X binding box labels to identify species, overcoming its limitation with hard to classify objects in images. Additionally, Phi-3.5 detects broader variables, such as vegetation type, and time of day, providing rich ecological and environmental context to YOLO's species detection output. When combined, this output is processed by the model's natural language system to answer complex queries, and retrieval-augmented generation (RAG) is employed to enrich responses with external information, like species weight and IUCN status (information that cannot be obtained through direct visual analysis). This information is used to automatically generate structured reports, providing biodiversity stakeholders with deeper insights into, for example, species abundance, distribution, animal behaviour, and habitat selection. Our approach delivers contextually rich narratives that aid in wildlife management decisions. By providing contextually rich insights, our approach not only reduces manual effort but also supports timely decision-making in conservation, potentially shifting efforts from reactive to proactive management.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。