用机器学习预测农药对蜜蜂毒性,发现现有模型不适用农业化学领域。
Evaluating machine learning models for predicting pesticide toxicity to honey bees
- 对比多种模型在蜜蜂毒性的预测表现
- 发现生物医学模型在农业数据上性能显著下降
- 强调需为农药研发专门构建数据与模型
小分子在生物医药、环境和农化领域具有关键作用,但各领域对理化性质和成功标准要求不同。尽管生物医药研究拥有丰富数据集和基准测试,农化领域的数据仍十分稀缺,尤其是针对物种特异性毒性的数据。本研究聚焦于目前最全面的实验验证农药对西方蜜蜂(Apis mellifera)毒性的数据集ApisTox。主要目标是评估多种机器学习方法在该毒性建模中的适用性,包括分子指纹、图核、图神经网络及预训练模型。与MoleculeNet基准中的药物数据集相比,ApisTox代表了一个独特的化学空间。在非医药数据集上的性能退化表明,仅基于生物医药数据训练的前沿算法泛化能力有限。研究强调需要更丰富的多样化数据集,并开发面向农化领域的专用模型。
原文摘要 · Abstract (English)
Small molecules play a critical role in the biomedical, environmental, and agrochemical domains, each with distinct physicochemical requirements and success criteria. Although biomedical research benefits from extensive datasets and established benchmarks, agrochemical data remain scarce, particularly with respect to species-specific toxicity. This work focuses on ApisTox, the most comprehensive dataset of experimentally validated chemical toxicity to the honey bee (\textit{Apis mellifera}), an ecologically vital pollinator. The primary goal of this study was to determine the suitability of diverse machine learning approaches for modeling such toxicity, including molecular fingerprints, graph kernels, and graph neural networks, as well as pretrained models. Comparative analysis with medicinal datasets from the MoleculeNet benchmark reveals that ApisTox represents a distinct chemical space. Performance degradation on non-medicinal datasets, such as \mbox{ApisTox}, demonstrates their limited generalizability of current state-of-the-art algorithms trained solely on biomedical data. Our study highlights the need for more diverse datasets and for targeted model development geared toward the agrochemical domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。