用大模型对话推荐科学数据,让研究者一键找到匹配的科研资源。
ScienceDB AI: An LLM-Driven Agentic Recommender System for Large-Scale Scientific Data Sharing Services

- 通过自然语言对话理解科研意图,提取实验要素进行精准推荐。
- 在超1000万数据集上验证,推荐准确率显著优于传统方法。
- 支持可追溯推荐,适合需要高效找数据的科研人员使用。
AI for Science 的快速发展凸显了科学数据集的重要性,推动了众多国家级数据共享平台的建立。然而,如何高效促进数据共享与利用仍面临挑战。科学数据蕴含复杂的领域知识,传统协同过滤推荐系统难以应对。大语言模型(LLM)为构建具备深度语义理解与个性化推荐能力的对话代理提供了新机遇。为此,我们提出 ScienceDB AI,一个基于 Science Data Bank(全球最大的科学数据共享平台之一)的 LLM 驱动型智能推荐系统。该系统通过自然语言对话与深度推理,精准推荐符合研究人员科研意图和动态需求的数据集。创新包括:科学意图感知器(Scientific Intention Perceptor),用于从复杂查询中提取结构化实验要素;结构化记忆压缩器(Structured Memory Compressor),有效管理多轮对话;以及可信检索增强生成框架(Trustworthy RAG),采用两阶段检索机制,并通过可引用的科学任务记录(Citable Scientific Task Record, CSTR)标识符提供可溯源的数据引用,提升推荐可信度与可复现性。基于超过 1000 万真实数据集的离线与在线实验表明,ScienceDB AI 具有显著有效性。据我们所知,这是首个专为大规模科学数据共享服务设计的 LLM 驱动型对话推荐系统。平台已公开访问:https://ai.scidb.cn/en。
原文摘要 · Abstract (English)
The rapid growth of AI for Science (AI4S) has underscored the significance of scientific datasets, leading to the establishment of numerous national scientific data centers and sharing platforms. Despite this progress, efficiently promoting dataset sharing and utilization for scientific research remains challenging. Scientific datasets contain intricate domain-specific knowledge and contexts, rendering traditional collaborative filtering-based recommenders inadequate. Recent advances in Large Language Models (LLMs) offer unprecedented opportunities to build conversational agents capable of deep semantic understanding and personalized recommendations. In response, we present ScienceDB AI, a novel LLM-driven agentic recommender system developed on Science Data Bank (ScienceDB), one of the largest global scientific data-sharing platforms. ScienceDB AI leverages natural language conversations and deep reasoning to accurately recommend datasets aligned with researchers' scientific intents and evolving requirements. The system introduces several innovations: a Scientific Intention Perceptor to extract structured experimental elements from complicated queries, a Structured Memory Compressor to manage multi-turn dialogues effectively, and a Trustworthy Retrieval-Augmented Generation (Trustworthy RAG) framework. The Trustworthy RAG employs a two-stage retrieval mechanism and provides citable dataset references via Citable Scientific Task Record (CSTR) identifiers, enhancing recommendation trustworthiness and reproducibility. Through extensive offline and online experiments using over 10 million real-world datasets, ScienceDB AI has demonstrated significant effectiveness. To our knowledge, ScienceDB AI is the first LLM-driven conversational recommender tailored explicitly for large-scale scientific dataset sharing services. The platform is publicly accessible at: https://ai.scidb.cn/en.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。