指导团队如何在数据整理阶段提升AI系统的可信度。
CaTE Data Curation for Trustworthy AI
- 梳理数据收集与处理流程,明确可信性标准。
- 提供可选路径与工具,覆盖不同场景需求。
- 适合数据科学家和AI系统开发者参考实践。
本报告为设计或开发AI系统团队提供实用指导,旨在促进开发过程中数据整理阶段的可信性。文中首先定义了数据、数据整理阶段及可信性的概念,随后列出一系列开发团队(尤其是数据科学家)可采取的步骤,以构建可信的AI系统。报告详细描述了核心步骤的序列,并追踪存在替代方案的并行路径。每一步均包含优势、局限、前提条件、预期结果以及相关开源工具实现。整体内容综合了学术文献中的数据整理方法与工具,目标是为读者提供一套多样且连贯的实践策略,以提升AI系统的可信度。
原文摘要 · Abstract (English)
This report provides practical guidance to teams designing or developing AI-enabled systems for how to promote trustworthiness during the data curation phase of development. In this report, the authors first define data, the data curation phase, and trustworthiness. We then describe a series of steps that the development team, especially data scientists, can take to build a trustworthy AI-enabled system. We enumerate the sequence of core steps and trace parallel paths where alternatives exist. The descriptions of these steps include strengths, weaknesses, preconditions, outcomes, and relevant open-source software tool implementations. In total, this report is a synthesis of data curation tools and approaches from relevant academic literature, and our goal is to equip readers with a diverse yet coherent set of practices for improving AI trustworthiness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。