将数据科学视为融合多维度复杂性的自然生态系统。
Data Science: a Natural Ecosystem
- 从数据生命周期出发,构建融合5D复杂性的数据科学框架。
- 提出计算与基础双分支的系统性划分,支持跨学科整合。
- 适合关注跨领域数据融合与智能系统设计的研究者。
本文从数据驱动的系统视角出发,提出一种称为‘本质数据科学’的自然生态系统,其核心挑战源于数据宇宙与五维复杂性(数据结构、领域、基数、因果关系、伦理)在数据生命周期各阶段的融合。数据代理根据特定目标执行任务,数据科学家作为数据代理逻辑组织的抽象实体存在。研究定义了特定学科驱动的数据科学,并由此衍生出涵盖各学科的泛数据科学生态系统。通过语义拆分,将本质数据科学划分为计算与基础两部分。该框架提供了一种通用、面向融合的架构,支持异构知识、代理与工作流的集成,适用于广泛学科与高影响力应用。
原文摘要 · Abstract (English)
This manuscript provides a systemic and data-centric view of what we term essential data science, as a natural ecosystem with challenges and missions stemming from the fusion of data universe with its multiple combinations of the 5D complexities (data structure, domain, cardinality, causality, and ethics) with the phases of the data life cycle. Data agents perform tasks driven by specific goals. The data scientist is an abstract entity that comes from the logical organization of data agents with their actions. Data scientists face challenges that are defined according to the missions. We define specific discipline-induced data science, which in turn allows for the definition of pan-data science, a natural ecosystem that integrates specific disciplines with the essential data science. We semantically split the essential data science into computational, and foundational. By formalizing this ecosystemic view, we contribute a general-purpose, fusion-oriented architecture for integrating heterogeneous knowledge, agents, and workflows-relevant to a wide range of disciplines and high-impact applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。