arXiv:2511.19752cs.CV2025-11

用可解释的原型网络,智能省去昂贵的基因检测。

What You See is (Usually) What You Get: Multimodal Prototype Networks that Abstain from Expensive Modalities

  • 基于多模态原型网络,自动判断何时无需基因数据
  • 图像+基因数据组合准确率与全量使用相当
  • 适合生态监测中需节省样本采集成本的场景

物种检测对生态系统健康评估和入侵物种识别至关重要,是指导保护行动的关键。多模态神经网络被广泛用于自动化物种识别,但存在两大缺陷:一是决策过程黑箱化,难以解释;二是获取基因数据成本高昂,常需捕获或杀死标本。本文通过扩展原型网络(ProtoPNets),提出一种可解释且成本感知的多模态方法。通过融合各模态原型并引入权重机制,动态决定预测对各模态的依赖程度。进一步设计策略识别无需昂贵基因信息即可自信分类的情况。实验表明,该方法能智能分配基因数据,仅在需要时使用,利用丰富图像数据完成清晰视觉分类,整体准确率与始终使用双模态的模型相当。

原文摘要 · Abstract (English)

Species detection is important for monitoring the health of ecosystems and identifying invasive species, serving a crucial role in guiding conservation efforts. Multimodal neural networks have seen increasing use for identifying species to help automate this task, but they have two major drawbacks. First, their black-box nature prevents the interpretability of their decision making process. Second, collecting genetic data is often expensive and requires invasive procedures, often necessitating researchers to capture or kill the target specimen. We address both of these problems by extending prototype networks (ProtoPNets), which are a popular and interpretable alternative to traditional neural networks, to the multimodal, cost-aware setting. We ensemble prototypes from each modality, using an associated weight to determine how much a given prediction relies on each modality. We further introduce methods to identify cases for which we do not need the expensive genetic information to make confident predictions. We demonstrate that our approach can intelligently allocate expensive genetic data for fine-grained distinctions while using abundant image data for clearer visual classifications and achieving comparable accuracy to models that consistently use both modalities.

物种识别原型网络多模态成本优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。