arXiv:2606.18302q-bio.OTcs.LG2026-06

首个孟加拉国本土鱼类蛋白序列数据集及高效识别模型,助力粮食安全与生态保护。

Protein-Based Fish Species Identification: Dataset, Models, and Insights from Native Bangladeshi Fish

论文配图:Protein-Based Fish Species Identification: Dataset, Models, and Insights from Native Bangladeshi Fish
图 1 · 摘自论文原文
  • 构建9种本土鱼的2845条高质量蛋白序列数据集,提出混合结构MotifCNN-Transformer+TA-PE。
  • 新模型达79.80%准确率,比大型模型ProtBERT快5倍、小42倍,支持无GPU推理。
  • 适用于资源受限地区部署,为渔业管理与生物多样性保护提供技术基础。

正确识别鱼类物种对孟加拉国的粮食安全、经济发展和气候韧性至关重要。蛋白质序列直接反映功能与进化约束,是物种鉴定与生物多样性监测的关键。然而,目前尚无针对孟加拉国本土鱼类的蛋白序列基准数据集。本研究首次构建了涵盖9种本土鱼类的2845条高质量蛋白序列数据集,并系统评估了七种架构在该任务上的表现,建立首个蛋白序列分类基线。提出一种新型混合架构MotifCNN-Transformer+TA-PE,实现79.80%准确率与0.80宏F1。微调后的ProtBERT(420M参数)达到最高83.04%准确率,但其性能提升相对于新模型统计不显著(p=0.1120)。尽管如此,新模型在6/9类别中表现更优,且速度更快5倍、体积小42倍、批量大16倍,支持无GPU推理,更适合农村等资源受限地区部署。研究还揭示了系统发育关系对序列相似性的影响,为南亚蛋白质依赖型经济体的渔业管理、食品认证与生物多样性保护提供路径。

原文摘要 · Abstract (English)

Correct identification of fish species is highly significant for food security, economic development, and climate resilience in Bangladesh. Protein sequences directly reflect functional and evolutionary constraints which are important for species authentication and biodiversity monitoring. Yet there exists no benchmark for native Bangladeshi fish species identification from protein sequence. In this study, we addressed this gap by introducing the first curated dataset for nine native Bangladeshi fish species of 2845 high quality protein sequences. We also established the first protein sequence classification baseline for this domain through a systematic benchmarking of seven architectural paradigms. Moreover, we propose a realistic deployable novel hybrid architecture of MotifCNN and Transformer with Terminal-Aware Positional-Encoding (MotifCNN-Transformer+TA-PE). Our novel architecture achieves 79.80% accuracy with macro-F1 of 0.80. The highest 83.04% accuracy is achieved by finetuned protein language model ProtBERT that has 420M parameters and requires dual 16GB GPUs for inference. According to McNemar's test, ProtBERT's 3.24% accuracy gain over our MotifCNN-Transformer+TA-PE is statistically insignificant (p = 0.1120). Our novel architecture beats it among six of the nine classes in per class identification. Also our MotifCNN-Transformer+TA-PE is approximately 5x faster, 42x smaller, and supports 16x larger batch size than ProtBERT and has GPU free inference, making it more practical for deployment in resources constrained areas such as rural Bangladesh. Beyond this, our foundational work shows effects of phylogenetic relationships on sequence similarity and establishes pathways for fisheries management, food authentication and biodiversity conservation in South Asia's protein dependent economy.

蛋白序列鱼类识别农业AI南亚应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。