用数据使用关系图识别可信数据,助力科研决策。
Exploring the Social Life of Data: Finding Data You Can Trust

- 构建数据使用图谱,追踪数据在科研中的流转路径。
- 通过使用历史判断数据可靠性,提升信任度。
- 适合科研人员、数据管理者及平台建设者参考。
人工智能正在改变科学探究的规模与节奏,模型可搜索、整合并推理远超个人研究者认知范围的数据。然而这一扩展带来了新挑战:在生成可信科学结果前,需找到适配问题、足够可靠且具备充分上下文支持负责任解读的数据。随着数据日益丰富,难题已从‘找数据’转变为‘找可信数据’。本文探索如何利用数据在科研中被使用时积累的社会与实证证据,类比社会信任网络,评估其适用性与可信度。具体提出将数据使用图谱作为新型科学数据基础设施,连接数据与发表论文、研究者、机构、主题、软件、模型、工作流及其他数据,揭示数据的‘社会生命’——谁曾依赖它、为哪些问题、以何种组合、采用何种方法、产生何种影响。这些关联将零散实践痕迹转化为数据使用描述符,补充传统元数据,辅助信任与适用性判断。核心观点并非流行即可信,而是可通过恰当背景历史建立信任。因此,使用证据须结合生产质量、来源、治理、语义清晰度与社区验证。可行性与价值通过在国家数据平台(NDP)内实现原型数据洞察发现服务得到验证。
原文摘要 · Abstract (English)
Artificial intelligence is changing the scale and tempo of scientific inquiry. Models can now search, integrate, and reason over data far beyond data repositories familiar to any individual researcher. Yet this expansion creates a prior problem: before a model can produce a trustworthy scientific result, it must locate data that are appropriate for the question, sufficiently reliable for the intended analysis, and accompanied by enough context to support responsible interpretation. As data becomes increasingly abundant, the challenge of finding data has been overcome by the challenge of finding data that you can trust. This paper explores how the social and empirical evidence that accumulates when data are used in research can be used, analogous to social trust networks, to determine fit for purpose and trust. Specifically, the paper explores data-usage graphs as a new layer of scientific data infrastructure. A data-usage graph connects datasets to the publications, people, institutions, topics, software, models, workflows, and other datasets through which they are produced and used. These connections reveal the {\it social life of data:} who has relied on a source, for which questions, in what combinations, with which methods, and with what observable impact. They can turn scattered traces of practice into data-usage descriptors that complement conventional metadata and support judgments of trust and fitness for purpose. The central claim is not that popularity establishes trust, but that this can be grown with appropriate contextual history. Usage evidence must therefore be combined with production quality, provenance, governance, semantic clarity, and community validation. The feasibility and value of data usage graphs is demonstrated by implementing the prototype data insights discovery service within the National Data Platform (NDP).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。