arXiv:2603.00084cs.DLcs.AI2026-03被引 1

让AI研究助手高效读取论文,自动提取结构化数据。

DeepXiv-SDK: An Agentic Data Interface for Scientific Literature

  • 将网页和PDF等杂乱文献转为可直接调用的标准化数据
  • 支持命令行、API和Python SDK等多种使用方式
  • 适合需要批量分析论文的研究者和AI开发人员

大语言模型代理正加速科学研究,但数据获取仍是瓶颈:代理缺乏便捷检索工具,且需处理网页和PDF等非结构化、面向人类的数据,导致消耗大量计算资源、效率低下且证据查找脆弱。为此,本文提出DeepXiv-SDK,一种三层次智能数据接口。第一层数据层将非结构化文献转换为标准化的JSON格式,提升可用性与可访问性;第二层服务层提供即用型数据访问工具与灵活代理使用方式(包括CLI、MCP、Python SDK);第三层应用层内置智能代理,整合服务层基础工具以应对复杂数据需求。DeepXiv-SDK目前已覆盖完整ArXiv论文库,并每日同步新发布内容,未来可扩展至PubMed Central、bioRxiv、medRxiv、chemRxiv等开放论文库。项目提供RESTful API、开源Python SDK及网页演示,展示深度搜索与研究工作流。注册后免费使用。

原文摘要 · Abstract (English)

LLM-agents are increasingly used to accelerate the progress of scientific research. Yet a persistent bottleneck is data access: agents not only lack readily available tools for retrieval, but also have to work with unstrcutured, human-centric data on the Internet, such as HTML web-pages and PDF files, leading to excessive token consumption, limit working efficiency, and brittle evidence look-up. This gap motivates the development of \textit{an agentic data interface}, which is designed to enable agents to access and utilize scientific literature in a more effective, efficient, and cost-aware manner. In this paper, we introduce DeepXiv-SDK, which offers a three-layer agentic data interface for scientific literature. 1) Data Layer, which transforms unstructured, human-centric data into normalized and structured representations in JSON format, improving data usability and enabling progressive accessibility of the data. 2) Service Layer, which presents readily available tools for data access and ad-hoc retrieval. It also enables a rich form of agent usage, including CLI, MCP, and Python SDK. 3) Application Layer, which creates a built-in agent, packaging basic tools from the service layer to support complex data access demands. DeepXiv-SDK currently supports the complete ArXiv corpus, and is synchronized daily to incorporate new releases. It is designed to extend to all common open-access corpora, such as PubMed Central, bioRxiv, medRxiv, and chemRxiv. We release RESTful APIs, an open-source Python SDK, and a web demo showcasing deep search and deep research workflows. DeepXiv-SDK is free to use with registration.

AI科研数据接口文献解析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。