整合超760万数据集,实现跨源可信发现与语义导航
SeDa: A Unified System for Dataset Discovery and Multi-Entity Augmented Semantic Exploration
- 统一框架融合200+平台数据,标准化元数据并构建可扩展标签图
- 支持主题检索与跨域关联,覆盖率达主流平台的1.8倍以上
- 基于机构、企业等多实体导航,适合科研与政策分析场景
开放数据平台与研究库的持续扩张导致数据集生态碎片化,给跨源数据发现与解释带来挑战。为此,我们提出SeDa——一个统一的数据集发现、语义标注与多实体增强探索框架。SeDa整合了来自200多个平台的超过760万条数据集,涵盖政府、学术与工业领域。该框架首先进行语义提取与标准化,以统一异构元数据表示;在此基础上,通过主题标签机制构建可扩展的标签图,支持主题检索与跨域关联;同时,嵌入溯源保障模块,在标注过程中持续验证数据来源并监控链接可用性,确保可靠性与可追溯性。此外,SeDa采用多实体增强导航策略,将数据集组织在站点、机构与企业构成的知识空间中,实现超越传统搜索的上下文感知与溯源探索。与ChatPD、Google Dataset Search等主流平台的对比实验表明,SeDa在覆盖率、时效性与可追溯性方面均表现更优。总体而言,SeDa为可信、语义丰富且全局可扩展的数据集探索奠定了基础。
原文摘要 · Abstract (English)
The continuous expansion of open data platforms and research repositories has led to a fragmented dataset ecosystem, posing significant challenges for cross-source data discovery and interpretation. To address these challenges, we introduce SeDa--a unified framework for dataset discovery, semantic annotation, and multi-entity augmented navigation. SeDa integrates more than 7.6 million datasets from over 200 platforms, spanning governmental, academic, and industrial domains. The framework first performs semantic extraction and standardization to harmonize heterogeneous metadata representations. On this basis, a topic-tagging mechanism constructs an extensible tag graph that supports thematic retrieval and cross-domain association, while a provenance assurance module embedded within the annotation process continuously validates dataset sources and monitors link availability to ensure reliability and traceability. Furthermore, SeDa employs a multi-entity augmented navigation strategy that organizes datasets within a knowledge space of sites, institutions, and enterprises, enabling contextual and provenance-aware exploration beyond traditional search paradigms. Comparative experiments with popular dataset search platforms, such as ChatPD and Google Dataset Search, demonstrate that SeDa achieves superior coverage, timeliness, and traceability. Taken together, SeDa establishes a foundation for trustworthy, semantically enriched, and globally scalable dataset exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。