用逼真图像编辑揭示生态监测中模型决策的关键特征。
Photorealistic Inpainting for Perturbation-based Explanations in Ecological Monitoring
- 通过修复式扰动生成逼真局部图像,保留场景上下文。
- 在冰川湾海豹识别任务中,显著降低模型置信度并提升可解释性。
- 适合生态学家和AI可信部署研究者使用。
生态监测正越来越多依赖视觉模型,但其预测过程不透明限制了信任与实地应用。本文提出一种基于修复的扰动解释方法,生成保持场景一致性的逼真局部图像编辑,用于揭示物种识别与性状归因任务中的细粒度形态线索。在经过微调的YOLOv9检测器上,利用Segment-Anything-Model优化的掩码,实施两种干预:(i) 对象移除/替换(如将海豹替换为冰、水或船只),(ii) 将原动物合成至新背景。通过重新评分扰动图像(翻转率、置信度下降)及专家评审生态合理性与可解释性进行评估。结果表明,该方法能精确定位诊断结构,避免传统扰动的删除伪影,并提供领域相关洞察,支持专家验证,增强生态领域AI部署的可信度。
原文摘要 · Abstract (English)
Ecological monitoring is increasingly automated by vision models, yet opaque predictions limit trust and field adoption. We present an inpainting-guided, perturbation-based explanation technique that produces photorealistic, mask-localized edits that preserve scene context. Unlike masking or blurring, these edits stay in-distribution and reveal which fine-grained morphological cues drive predictions in tasks such as species recognition and trait attribution. We demonstrate the approach on a YOLOv9 detector fine-tuned for harbor seal detection in Glacier Bay drone imagery, using Segment-Anything-Model-refined masks to support two interventions: (i) object removal/replacement (e.g., replacing seals with plausible ice/water or boats) and (ii) background replacement with original animals composited onto new scenes. Explanations are assessed by re-scoring perturbed images (flip rate, confidence drop) and by expert review for ecological plausibility and interpretability. The resulting explanations localize diagnostic structures, avoid deletion artifacts common to traditional perturbations, and yield domain-relevant insights that support expert validation and more trustworthy deployment of AI in ecology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。