arXiv:2603.04380cs.CVcs.CL2026-03中稿 · WACV 2026被引 1

用中间奖励训练模型分步推理物种分类,准确率超人类且过程可解释。

TaxonRL: Reinforcement Learning with Intermediate Rewards for Interpretable Fine-Grained Visual Reasoning

  • 分层级预测物种、属、科特征,中间奖励引导逐步推理
  • 在Birds-to-Words上达91.7%准确率,超过人类77.3%
  • 结果可追溯,适合需要透明决策的生物分类场景

传统视觉-语言模型在区分同属或同科内视觉相似物种时表现不佳。我们提出TaxonRL,一种基于组相对策略优化与中间奖励的强化学习方法,将推理过程分解为层级化的分类判断。该方法激励模型在最终分类前显式分析物种级、属级和科级特征,不仅提升准确率,还生成可验证的决策路径。在Birds-to-Words数据集上,TaxonRL平均准确率达91.7%,超过人类表现(77.3%),并生成可解释的推理轨迹。我们进一步验证了其在灵长类与海洋物种识别中的强跨域泛化能力。结果表明,结构化层级推理是细粒度视觉区分的有效且可迁移的框架。

原文摘要 · Abstract (English)

Traditional vision-language models struggle with contrastive fine-grained taxonomic reasoning, particularly when distinguishing between visually similar species within the same genus or family. We introduce TaxonRL, a reinforcement learning approach using Group Relative Policy Optimization with intermediate rewards that decomposes the reasoning process into hierarchical taxonomic predictions. Our method incentivizes models to explicitly reason about species-level, genus-level, and family-level features before making final classifications. This structured approach is designed not only to boost accuracy but also to yield a transparent, verifiable decision-making process. On the challenging Birds-to-Words dataset, TaxonRL achieves 91.7\% average accuracy, exceeding human performance (77.3\%) while generating interpretable reasoning traces. We demonstrate strong cross-domain generalization, showing substantial gains in primate and marine species verification. Our results establish that enforcing structured, hierarchical reasoning provides a powerful and transferable framework for fine-grained visual discrimination.

细粒度分类强化学习可解释性视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。