首个针对分子毒性修复的基准测试,评估大模型改写有毒分子的能力。
Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?
- 构建涵盖660种毒物的标准化数据集,支持11类修复任务。
- 43个主流多模态大模型在毒性降低上表现有限但初现结构编辑能力。
- 适合药物研发、AI制药及毒性预测方向的研究者参考。
毒性仍是新药研发早期失败的主要原因。尽管分子设计与性质预测取得进展,但分子毒性修复——生成结构有效且毒性更低的替代分子——尚未被系统定义或评测。为此,我们提出ToxiMol,首个面向通用多模态大模型(MLLMs)的分子毒性修复基准任务。构建覆盖11项主任务、660个代表性有毒分子的标准数据集,涵盖多样机制与粒度。设计基于专家毒理知识的机制感知、任务自适应提示标注流程。同时提出自动化评估框架ToxiEval,集成毒性终点预测、合成可及性、类药性与结构相似性,形成高通量修复效果评估链。系统评估43个主流通用型MLLMs,并开展多重消融实验分析评估指标、候选多样性与失败归因等关键问题。结果表明,当前MLLMs在此任务仍面临显著挑战,但在毒性理解、语义约束遵守与结构感知编辑方面已初现潜力。
原文摘要 · Abstract (English)
Toxicity remains a leading cause of early-stage drug development failure. Despite advances in molecular design and property prediction, the task of molecular toxicity repair, generating structurally valid molecular alternatives with reduced toxicity, has not yet been systematically defined or benchmarked. To fill this gap, we introduce ToxiMol, the first benchmark task for general-purpose Multimodal Large Language Models (MLLMs) focused on molecular toxicity repair. We construct a standardized dataset covering 11 primary tasks and 660 representative toxic molecules spanning diverse mechanisms and granularities. We design a prompt annotation pipeline with mechanism-aware and task-adaptive capabilities, informed by expert toxicological knowledge. In parallel, we propose an automated evaluation framework, ToxiEval, which integrates toxicity endpoint prediction, synthetic accessibility, drug-likeness, and structural similarity into a high-throughput evaluation chain for repair success. We systematically assess 43 mainstream general-purpose MLLMs and conduct multiple ablation studies to analyze key issues, including evaluation metrics, candidate diversity, and failure attribution. Experimental results show that although current MLLMs still face significant challenges on this task, they begin to demonstrate promising capabilities in toxicity understanding, semantic constraint adherence, and structure-aware editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。