arXiv:2604.26645cs.AIcs.LG2026-04KDD

构建智能系统评估科学数据的AI可用性,解决跨领域数据质量难量化问题。

SciHorizon-DataEVA: An Agentic System for AI-Readiness Evaluation of Heterogeneous Scientific Data

论文配图:SciHorizon-DataEVA: An Agentic System for AI-Readiness Evaluation of Heterogeneous Scientific Data
图 1 · 摘自论文原文
  • 提出四维评估框架Sci-TQA2,涵盖治理可信、数据质量、模型适配与科学适应性。
  • 开发多智能体系统动态生成评估方案,在多领域数据上实现高效可靠测评。
  • 适合科研人员和数据工程师评估数据是否适合用于机器学习建模。

人工智能推动科学发现,但其效果受限于科学数据的AI就绪程度。目前尚无可扩展的系统化评估机制。本文提出SciHorizon-DataEVA,一个新型智能系统,用于异构科学数据的可扩展评估。我们引入Sci-TQA2原则,将AI就绪性划分为四个互补维度:治理可信度、数据质量、模型兼容性和科学适应性。每个维度分解为可测量的原子单元,支持细粒度评估。通过层次化多智能体架构,Sci-TQA2-Eval基于数据特征、任务适用性及文献信号,动态生成评估规范,并以工具为中心的自适应机制执行,具备验证与自我修正能力。在多领域科学数据集上的实验表明,该系统具备有效性和普适性。

原文摘要 · Abstract (English)

AI-for-Science (AI4Science) is increasingly transforming scientific discovery by embedding machine learning models into prediction, simulation, and hypothesis generation workflows across domains. However, the effectiveness of these models is fundamentally constrained by the AI-readiness of scientific data, for which no scalable and systematic evaluation mechanism currently exists. In this work, we propose SciHorizon-DataEVA, a novel agentic system to scalable AI-readiness evaluation of heterogeneous scientific data. At the evaluation-criteria level, we introduce the Sci-TQA2 principles, which organize AI-readiness into four complementary dimensions: Governance Trustworthiness, Data Quality, AI Compatibility, and Scientific Adaptability. Each dimension is decomposed into measurable atomic elements that enable fine-grained and executable assessment. To operationalize these principles at scale, we develop Sci-TQA2-Eval, a hierarchical multi-agent evaluation approach orchestrated through a directed, cyclic workflow. Our Sci-TQA2-Eval dynamically constructs dataset-aware evaluation specifications by combining lightweight dataset profiling, applicability-aware metric activation, and knowledge-augmented planning grounded in domain constraints and dataset-paper signals. These specifications are executed through an adaptive, tool-centric evaluation mechanism with built-in verification and self-correction, enabling scalable and reliable assessment across heterogeneous scientific data. Extensive experiments on scientific datasets spanning multiple domains demonstrate the effectiveness and generality of SciHorizon-DataEVA for principled AI-readiness evaluation.

AI就绪性科学数据多智能体评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。