AI自动生成生物信息学论文,确保每项结论有文献支撑且实验结果真实。
Prompt-to-Paper: Agentic AI System for Bioinformatics

- 用可验证文献库+引用扩展,让每个论点都有据可查。
- 自动执行计算实验,生成真实数值而非伪造结果。
- 八维评分体系+迭代优化,适合科研人员投稿前自查。
尽管大语言模型已实现端到端自动化论文生成,现有系统仍存在三大缺陷:(i) 生成论断缺乏可验证的文献依据,(ii) 实验结果常为虚构而非真实执行,(iii) 缺乏标准化、多维度的质量评估框架。我们提出 Prompt-to-Paper,一种多智能体系统,通过三项集成创新解决评估空白。首先,采用确定性检索增强生成流程,结合章节感知相关性评分与雪球式引用扩展,将每项论断锚定在60–100篇可验证论文构成的语料库中。其次,自主代码代理执行真实计算生物学实验,以实际数值替代合成输出。第三,基于已发表论文的近似参考统计并加入显式幻觉惩罚,构建八维自动化质量评分器,实现标准化、可复现的评估。质量驱动的改进循环包含上下文丰富的修订器,每轮引导至三种研究人员操作之一,并每十轮触发深度研究循环,重新运行实验并从更强输出重写论文。我们在五个生物信息学案例上验证该系统,所有案例均生成符合投稿格式的PDF,无超出范围引用。改进循环使论文质量平均提升17.96分(满分100,最高达26.04)。部分外部人工评审显示,五篇论文平均得分7.0/10。完整论文生成成本约为每篇0.31美元。
原文摘要 · Abstract (English)
While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three critical deficiencies: (i) generated claims are not deterministically grounded in verifiable literature, (ii) experimental results are frequently fabricated rather than executed, and (iii) there exists no standardized, multi-dimensional framework to assess whether AI-generated manuscripts meet the quality and rigor required for real-world publication. We present Prompt-to-Paper, a multi-agent framework that directly addresses this evaluation gap through three integrated innovations. First, a deterministic retrieval-augmented generation pipeline with section-aware relevance scoring and snowball citation expansion grounds every claim in a verifiable corpus of 60--100 papers. Second, an autonomous coding agent executes real computational biology experiments replacing synthetic outputs with genuine numerical results. Third, an eight-dimensional automated quality scorer, benchmarked with approximate reference statistics from published papers and augmented with explicit hallucination penalties, provides standardized, reproducible quality assessments. The quality-driven improvement loop uses a context-rich reviser that routes each iteration to one of three researcher actions and fires a deep research cycle every ten iterations to re-run experiments and re-manuscript from stronger outputs. We validate the system on five bioinformatics case studies; all five cases compiled submission-formatted PDFs with zero out-of-range citations. The improvement loop raises manuscript quality by an average of +17.96 points on a 0--100 scale (maximum +26.04. As partial external checks, a human reviewer scored the five manuscripts at an average of 7.0 out of 10. Complete manuscripts are produced at approximately 0.31 USD per paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。