重建原始Tox21数据集排行榜,验证十年来药物毒性预测是否真正进步。
Measuring AI Progress in Drug Discovery: A Reproducible Leaderboard for the Tox21 Challenge
- 基于原始Tox21数据构建可复现的排行榜,杜绝数据篡改影响结果
- 发现2015年冠军DeepTox和2017年自归一化网络仍位居前列
- 公开模型接口,供研究者直接调用验证,推动公平比较
深度学习自2010年代初兴起后,在计算机视觉与自然语言处理等领域取得突破,并深刻影响生物医学研究。2015年,深度神经网络在Tox21数据挑战赛中超越传统方法,成为药物发现领域的关键转折点,推动制药行业广泛采用深度学习技术。此后,Tox21数据集被纳入MoleculeNet、Open Graph Benchmark等主流基准,但过程中数据被修改,标签被填补或生成,导致各研究间结果不可比。为此,本文在Hugging Face上发布基于原始Tox21数据集的可复现排行榜,包含基线与代表性模型。当前结果显示,2015年冠军方法(基于集成的DeepTox)及2017年提出的描述符型自归一化神经网络仍表现优异,位列顶尖,表明过去十年毒性预测进展尚不明确。本工作同时提供所有基线与评估模型的标准化API接口,可通过Hugging Face Spaces在线推理。
原文摘要 · Abstract (English)
Deep learning's rise since the early 2010s has transformed fields like computer vision and natural language processing and strongly influenced biomedical research. For drug discovery specifically, a key inflection - akin to vision's "ImageNet moment" - arrived in 2015, when deep neural networks surpassed traditional approaches on the Tox21 Data Challenge. This milestone accelerated the adoption of deep learning across the pharmaceutical industry, and today most major companies have integrated these methods into their research pipelines. After the Tox21 Challenge concluded, its dataset was included in several established benchmarks, such as MoleculeNet and the Open Graph Benchmark. However, during these integrations, the dataset was altered and labels were imputed or manufactured, resulting in a loss of comparability across studies. Consequently, the extent to which bioactivity and toxicity prediction methods have improved over the past decade remains unclear. To this end, we introduce a reproducible leaderboard, hosted on Hugging Face with the original Tox21 Challenge dataset, together with a set of baseline and representative methods. The current version of the leaderboard indicates that the original Tox21 winner - the ensemble-based DeepTox method - and the descriptor-based self-normalizing neural networks introduced in 2017, continue to perform competitively and rank among the top methods for toxicity prediction, leaving it unclear whether substantial progress in toxicity prediction has been achieved over the past decade. As part of this work, we make all baselines and evaluated models publicly accessible for inference via standardized API calls to Hugging Face Spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。