AI科学家直接读取原始多模态数据,全程自主完成科研全流程。
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

- 构建感知层+三智能体闭环,从原始数据直接生成研究假设
- 36个真实案例全部完成从数据到论文的全流程,平均得分6.3
- 直接感知原始数据比仅用预计算特征提升85%判断胜率
近年来基础模型的发展使AI科学家能够自动化涵盖假说生成、代码执行到论文撰写的完整研究流程。然而,仅覆盖工作流并不足以获取科学发现所依赖的全部证据。现有系统通常仅对文本、代码、标签或预计算摘要进行推理,导致空间、时间、跨通道和过程性关系无法被智能体利用。我们提出OmniScientist,一个端到端的全模态、跨学科AI科学家,可直接从异构原始证据中开展多学科研究。系统包含感知层与三个自治智能体(构思、实验、撰写),在确定性流程中实现观测驱动的研究问题、实验决策与最终结论的动态演化。通过代码执行的思路、严谨性与主张检查,系统确保新颖性筛选、统计有效性、执行溯源与数值可追溯性。我们在36个真实数据案例上评估该系统,涵盖5个学科族、4类科学证据,包括图像、信号、音频、视频、三维结构、轨迹、表格、公式与图表等模态。系统在所有案例中均完成从原始数据到成稿论文的全过程,采用参考推理主干时平均论文得分为6.3。与仅接收预计算标量特征的盲测版本相比,直接感知原始数据在7项评估维度上全面领先,且在85%的两两对比中胜出。结果表明,全生命周期的感知对于基于证据的科学发现至关重要,并为构建通用型AI科学家提供了可行路径。
原文摘要 · Abstract (English)
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。