为科研数据设计可被智能体使用的标准化技能接口。
Scientific Data Skills: Enabling Agent-Ready Scientific Data Services at Scale

- 将数据描述、使用方法等知识封装为可复用的智能体技能
- 在六大学科领域构建技能库,支持精准发现与解释数据
- 适合需要自动化处理科研数据的AI研究者和开发者
科学数据正被越来越多的AI智能体使用,但现有数据集表示对自主发现、解读和调用的支持有限,这源于科学数据在异构存储库中的分散性,以及主要面向人类的设计。为此,我们提出科学数据技能(SciDSK),一种面向智能体的数据表示形式,将数据特定知识和操作指引打包为可复用的智能体技能。SciDSK整合了数据描述、科学背景、文件结构、使用流程、质量检查与溯源信息,同时保持数据原样存放于原始仓库。我们定义了结构化规范,并建立系统化构建管道,确保每个SciDSK基于权威数据记录和辅助材料。进一步构建了科学数据技能库,作为统一平台,在六大学科领域发布SciDSK资源,支持包访问、持久标识与溯源。通过数据发现检索基准与受控案例评估,结果表明SciDSK显著提升智能体驱动的数据发现能力,并提供更精确、可操作的数据解读支持。这些发现验证了以智能体就绪方式组织数据知识的价值。
原文摘要 · Abstract (English)
Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for autonomous discovery, interpretation, and invocation. This limitation stems from the fragmentation of scientific data across heterogeneous repositories and from dataset representations designed primarily for human use. To address this limitation, we introduce the Scientific Data Skill (SciDSK), an agent-ready representation that packages dataset-specific knowledge and operational guidance as a reusable agent skill. A SciDSK integrates dataset descriptions, scientific context, file organization, usage procedures, quality checks, and provenance information while retaining the underlying data in its original repository. We define a structured SciDSK specification and develop a systematic construction pipeline that grounds each SciDSK in authoritative dataset records and associated supporting materials. We further establish the Scientific Data Skill Bank, a unified platform that publishes SciDSK resources across six scientific disciplines and supports package access, persistent identification, and traceability to source datasets. We evaluate SciDSK through a retrieval benchmark for dataset discovery and controlled cases for dataset interpretation. The results show that SciDSK improves agent-driven dataset discovery and provides more precise and actionable support for dataset interpretation. These findings support the value of organizing dataset-specific knowledge in an agent-ready representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。