arXiv:2509.09738cs.AIq-bio.QM2025-09被引 1

AI助手将药物申报文档撰写时间减少97%,质量达标但需人工优化。

Human-AI Collaboration Increases Efficiency in Regulatory Writing

  • 用大语言模型自动生成临床前报告,大幅缩短初稿时间。
  • 18870页文档初稿仅需3.7小时,质量评分超69%,无严重错误。
  • 适合药企研发人员和监管写作者参考,助力AI协同提效。

背景:新药临床试验申请(IND)准备耗时且依赖专业经验,制约早期临床开发。目标:评估大型语言模型平台(AutoIND)能否在保持文档质量的前提下缩短首次起草时间。方法:直接记录AutoIND生成IND非临床总结报告(eCTD模块2.6.2、2.6.4、2.6.6)的起草时间;以美国FDA已批准的报告为基准,估算资深注册写作人员(≥6年经验)的手动撰写时间作为行业标准。质量由盲评专家依据正确性、完整性、简洁性、一致性、清晰度、冗余度和重点突出等七项指标评分(0–3分,归一化为百分比)。关键监管错误定义为可能改变监管判断的误报或遗漏(如错误的NOAEL值、遗漏强制性GLP剂量-制剂分析)。结果:AutoIND使初稿时间降低约97%(从约100小时降至IND-1的3.7小时,覆盖18,870页/61份报告;IND-2为2.6小时,11,425页/58份报告)。质量评分分别为69.6%和77.9%。未发现关键监管错误,但存在重点突出、简洁性和清晰度方面的不足。结论:AutoIND可显著加速IND起草,但专家仍需介入以达到提交可用水平。识别出的系统性缺陷为模型改进提供了明确方向。

原文摘要 · Abstract (English)

Background: Investigational New Drug (IND) application preparation is time-intensive and expertise-dependent, slowing early clinical development. Objective: To evaluate whether a large language model (LLM) platform (AutoIND) can reduce first-draft composition time while maintaining document quality in regulatory submissions. Methods: Drafting times for IND nonclinical written summaries (eCTD modules 2.6.2, 2.6.4, 2.6.6) generated by AutoIND were directly recorded. For comparison, manual drafting times for IND summaries previously cleared by the U.S. FDA were estimated from the experience of regulatory writers ($\geq$6 years) and used as industry-standard benchmarks. Quality was assessed by a blinded regulatory writing assessor using seven pre-specified categories: correctness, completeness, conciseness, consistency, clarity, redundancy, and emphasis. Each sub-criterion was scored 0-3 and normalized to a percentage. A critical regulatory error was defined as any misrepresentation or omission likely to alter regulatory interpretation (e.g., incorrect NOAEL, omission of mandatory GLP dose-formulation analysis). Results: AutoIND reduced initial drafting time by $\sim$97% (from $\sim$100 h to 3.7 h for 18,870 pages/61 reports in IND-1; and to 2.6 h for 11,425 pages/58 reports in IND-2). Quality scores were 69.6\% and 77.9\% for IND-1 and IND-2. No critical regulatory errors were detected, but deficiencies in emphasis, conciseness, and clarity were noted. Conclusions: AutoIND can dramatically accelerate IND drafting, but expert regulatory writers remain essential to mature outputs to submission-ready quality. Systematic deficiencies identified provide a roadmap for targeted model improvements.

AI协作药物研发大模型应用监管写作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。