DataJoint 2.0用关系模型统一科学工作流的数据与计算,让人类和智能体协作更可靠。
DataJoint 2.0: A Computational Substrate for Agentic Scientific Workflows

- 用表格和行表示工作流步骤与产物,外键定义执行顺序。
- 支持对象存储、语义匹配和领域格式扩展,提升数据一致性。
- 适合需要可复现、可验证的自动化科研团队使用。
操作严谨性决定了人机协作是否成功。科学数据流程需要类似DevOps的SciOps体系,但现有方法常因系统割裂导致溯源断裂且无事务保障。DataJoint 2.0通过关系工作流模型解决此问题:表代表工作流步骤,行代表产物,外键规定执行顺序。模式不仅定义数据存在形式,还明确其生成方式——形成一个统一的可查询、可强制、机器可读的系统。四项技术突破延伸该基础:融合关系元数据与可扩展对象存储的增强型模式、基于属性溯源的语义匹配防止错误连接、支持领域特定格式的可扩展类型系统,以及与外部编排工具兼容的分布式任务协调机制。通过统一数据结构、数据本身与计算转换,DataJoint构建了支持SciOps的底层平台,使智能体可在不破坏数据完整性的前提下参与科学工作流。
原文摘要 · Abstract (English)
Operational rigor determines whether human-agent collaboration succeeds or fails. Scientific data pipelines need the equivalent of DevOps -- SciOps -- yet common approaches fragment provenance across disconnected systems without transactional guarantees. DataJoint 2.0 addresses this gap through the relational workflow model: tables represent workflow steps, rows represent artifacts, foreign keys prescribe execution order. The schema specifies not only what data exists but how it is derived -- a single formal system where data structure, computational dependencies, and integrity constraints are all queryable, enforceable, and machine-readable. Four technical innovations extend this foundation: object-augmented schemas integrating relational metadata with scalable object storage, semantic matching using attribute lineage to prevent erroneous joins, an extensible type system for domain-specific formats, and distributed job coordination designed for composability with external orchestration. By unifying data structure, data, and computational transformations, DataJoint creates a substrate for SciOps where agents can participate in scientific workflows without risking data corruption.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。