arXiv:2603.25573cs.CVcs.LG2026-03中稿 · ICLR

通过编码生物分类层级结构,提升多模态物种识别的准确性和鲁棒性。

Hierarchy-Guided Multimodal Representation Learning for Taxonomic Inference

  • 引入层级信息正则化,使嵌入表示符合分类层级关系。
  • 在部分或损坏的DNA数据下,分类准确率提升超14%。
  • 支持单模态与多模态推理,适合实际生态监测场景。

从大规模野外数据中精准识别生物多样性是生态学、保护生物学和环境监测的基础问题。核心任务是基于不完整输入(如标本图像、DNA条形码或两者结合)进行分类预测,即推断物种的纲、目、科、属或种。现有多模态方法通常将分类视为平坦标签空间,未能捕捉生物分类的层级结构,导致在噪声或模态缺失时表现不佳。本文提出两种端到端的层次感知多模态学习方法:CLiBD-HiR通过层级信息正则化(HiR)塑造跨分类层级的嵌入几何结构,生成结构化且抗噪的表示;CLiBD-HiR-Fuse进一步训练轻量级融合预测器,支持仅图像、仅DNA或联合推理,对模态损坏具有鲁棒性。在多个大规模生物多样性基准上,相比强基线方法,分类准确率提升超过14%,尤其在部分或受损的DNA条件下增益显著。结果表明,显式编码生物层级结构并结合灵活融合,是构建实用生物多样性基础模型的关键。

原文摘要 · Abstract (English)

Accurate biodiversity identification from large-scale field data is a foundational problem with direct impact on ecology, conservation, and environmental monitoring. In practice, the core task is taxonomic prediction - inferring order, family, genus, or species from imperfect inputs such as specimen images, DNA barcodes, or both. Existing multimodal methods often treat taxonomy as a flat label space and therefore fail to encode the hierarchical structure of biological classification, which is critical for robustness under noise and missing modalities. We present two end-to-end variants for hierarchy-aware multimodal learning: CLiBD-HiR, which introduces Hierarchical Information Regularization (HiR) to shape embedding geometry across taxonomic levels, yielding structured and noise-robust representations; and CLiBD-HiR-Fuse, which additionally trains a lightweight fusion predictor that supports image-only, DNA-only, or joint inference and is resilient to modality corruption. Across large-scale biodiversity benchmarks, our approach improves taxonomic classification accuracy by over 14 percent compared to strong multimodal baselines, with particularly large gains under partial and corrupted DNA conditions. These results highlight that explicitly encoding biological hierarchy, together with flexible fusion, is key for practical biodiversity foundation models.

多模态分类生物多样性层级结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。