用AI自动标注生物图像的形态特征,让机器理解昆虫细节。
Automatic Image-Level Morphological Trait Annotation for Organismal Images
- 用稀疏自编码器提取图像中专注特定部位的神经元
- 构建8万条标注数据集,覆盖1.9万张昆虫图像
- 适合生态学与计算机视觉交叉研究者使用
形态特征是生物体的物理特性,为理解生物与环境互动提供关键线索。但当前提取这些特征仍依赖人工、耗时费力,限制了其在大规模生态研究中的应用。主要瓶颈在于缺乏高质量的图像-特征标注数据集。本文证明,基于基础模型特征训练的稀疏自编码器可生成语义单一、空间定位精准的神经元,稳定激活于有意义的形态区域。利用此特性,我们提出一种形态特征标注流程:先定位显著区域,再通过视觉语言提示生成可解释的特征描述。基于该方法,我们构建了Bioscan-Traits数据集,包含80,000条标注,涵盖来自BIOSCAN-5M的19,000张昆虫图像。人工评估确认生成的形态描述具有生物学合理性。通过全面消融实验,系统分析关键设计选择对描述质量的影响。相比昂贵的人工标注,本方法以模块化方式实现可扩展的生物意义监督,推动基础模型注入生态相关性,支持大规模形态分析,弥合生态实用性与机器学习之间的鸿沟。
原文摘要 · Abstract (English)
Morphological traits are physical characteristics of biological organisms that provide vital clues on how organisms interact with their environment. Yet extracting these traits remains a slow, expert-driven process, limiting their use in large-scale ecological studies. A major bottleneck is the absence of high-quality datasets linking biological images to trait-level annotations. In this work, we demonstrate that sparse autoencoders trained on foundation-model features yield monosemantic, spatially grounded neurons that consistently activate on meaningful morphological parts. Leveraging this property, we introduce a trait annotation pipeline that localizes salient regions and uses vision-language prompting to generate interpretable trait descriptions. Using this approach, we construct Bioscan-Traits, a dataset of 80K trait annotations spanning 19K insect images from BIOSCAN-5M. Human evaluation confirms the biological plausibility of the generated morphological descriptions. We assess design sensitivity through a comprehensive ablation study, systematically varying key design choices and measuring their impact on the quality of the resulting trait descriptions. By annotating traits with a modular pipeline rather than prohibitively expensive manual efforts, we offer a scalable way to inject biologically meaningful supervision into foundation models, enable large-scale morphological analyses, and bridge the gap between ecological relevance and machine-learning practicality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。