用分类树结构生成稀有物种图像,兼顾特异细节与共性特征。
TreeAdapter: Hierarchical Taxonomy-Guided Adapter Composition for Fine-Grained Species Image Generation

- 在分类树每个节点加轻量适配器,叶节点抓特有特征,内部节点存共性语义。
- 两阶段训练使祖先适配器只学后代未覆盖的视觉残差,提升生成精度。
- 适合生物多样性研究、濒危物种图像合成,对细粒度图像生成效果显著。
尽管通用文本到图像模型在开放领域表现优异,但在生成稀有生物物种图像等特定下游任务中性能明显下降。由于长尾分布问题,通用模型难以捕捉细微的细粒度特征;而针对每种物种的微调方法又过度隔离个体,忽略近缘类群间的共享视觉特征。为此,我们提出TreeAdapter,一种显式利用分类层级数据的新框架。不采用单一模型或独立的物种级模块,而是将轻量适配器附加到分类树的每一个节点上:叶节点适配器捕获物种特异性视觉特征,内部节点适配器封装其子代类群的共享语义。我们设计了两阶段训练范式,使祖先适配器仅优化其后代无法解释的视觉残差。该架构与训练策略使模型充分挖掘层级信息,确保每种物种视觉特征的准确生成。在三个大规模生物多样性基准上的大量实验表明,TreeAdapter在细粒度图像生成质量上达到当前最优水平,优于通用和领域专用基线模型。
原文摘要 · Abstract (English)
Although general text-to-image models excel in open-domain generation, their performance degrades significantly in specialized downstream domains, particularly when generating images of rare biological species. Hindered by long-tailed distributions, general models struggle to capture subtle fine-grained details, while per-species fine-tuning methods over-isolate individual species and consequently ignore the shared visual features among closely related taxa. To address this, we propose TreeAdapter, a novel framework that explicitly leverages hierarchical taxonomic data. Rather than using a monolithic model or independent per-species modules, TreeAdapter attaches lightweight adapters to every node of the taxonomic tree. Specifically, leaf-node adapters capture species-specific visual traits, while internal-node adapters encapsulate shared semantics among descendant taxa. We introduce a two-stage training paradigm where ancestor adapters are optimized to model only the residual visual features unexplained by their descendants. This model architecture and training paradigm enable the model to fully leverage hierarchical information, ensuring the accurate generation of visual features for each species. Extensive experiments across three large-scale biodiversity benchmarks demonstrate that TreeAdapter achieves state-of-the-art fine-grained generation quality, outperforming both general-purpose and domain-specific baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。