用生物分类模型提升物种级图像生成精度。
TaxaAdapter: Vision Taxonomy Models are Key to Fine-grained Image Generation over the Tree of Life
- 引入视觉分类模型嵌入,指导细粒度物种生成。
- 在多种物种上显著提升形态准确性和身份识别率。
- 适合生物多样性研究与少样本物种生成场景。
准确生成生命之树中超过1000万种物种的图像极具挑战,因许多物种仅通过细微视觉特征区分。尽管文本到图像合成取得进展,现有模型仍难以捕捉定义物种身份的细粒度视觉线索,即使输出图像逼真。为此,我们提出TaxaAdapter,一种轻量级方法,通过将视觉分类模型(如BioCLIP)的嵌入注入冻结的文本到图像扩散模型,以提升物种级保真度,同时保持对姿态、风格和背景等属性的灵活控制。大量实验表明,TaxaAdapter在强基线基础上持续提升形态保真度与物种身份准确性,且架构更简洁、训练更稳定。为更好评估,我们引入基于多模态大语言模型的指标,从生成与真实图像中总结性状描述,提供更可解释的形态一致性度量。此外,TaxaAdapter展现出强泛化能力,可实现少样本物种(仅需少量训练图像)甚至训练未见物种的合成。结果表明,视觉分类模型是实现可扩展细粒度物种生成的关键。
原文摘要 · Abstract (English)
Accurately generating images across the Tree of Life is difficult: there are over 10M distinct species on Earth, many of which differ only by subtle visual traits. Despite the remarkable progress in text-to-image synthesis, existing models often fail to capture the fine-grained visual cues that define species identity, even when their outputs appear photo-realistic. To this end, we propose TaxaAdapter, a simple and lightweight approach that incorporates Vision Taxonomy Models (VTMs) such as BioCLIP to guide fine-grained species generation. Our method injects VTM embeddings into a frozen text-to-image diffusion model, improving species-level fidelity while preserving flexible text control over attributes such as pose, style, and background. Extensive experiments demonstrate that TaxaAdapter consistently improves morphology fidelity and species-identity accuracy over strong baselines, with a cleaner architecture and training recipe. To better evaluate these improvements, we also introduce a multimodal Large Language Model-based metric that summarizes trait-level descriptions from generated and real images, providing a more interpretable measure of morphological consistency. Beyond this, we observe that TaxaAdapter exhibits strong generalization capabilities, enabling species synthesis in challenging regimes such as few-shot species with only a handful of training images and even species unseen during training. Overall, our results highlight that VTMs are a key ingredient for scalable, fine-grained species generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。