arXiv:2502.19567cs.CRcs.AI2025-02被引 18

Atlas让机器学习全生命周期可追溯,防数据污染还保隐私。

Atlas: A Framework for ML Lifecycle Provenance & Transparency

  • 用开放标准记录模型全链路来源与真实性证据。
  • 结合可信硬件与透明日志,保障元数据完整且不泄露数据。
  • 适合需合规、防供应链攻击的AI系统开发者。

开源机器学习数据集和模型的广泛应用,使当前人工智能应用面临数据投毒和供应链攻击等重大风险。在日益增长的监管压力下,模型供应商需在提升透明度与保护数据及知识产权之间取得平衡。本文提出Atlas框架,实现可验证的机器学习全流程可追溯性。Atlas利用数据与软件供应链溯源的开放规范,收集模型制品真实性的可验证记录及端到端的血缘元数据。通过融合可信硬件与透明日志,Atlas增强元数据完整性,保护数据机密性,并限制训练到部署过程中的未授权访问。我们构建了原型系统,集成多个开源工具,通过两个案例研究评估其实际可行性。

原文摘要 · Abstract (English)

The rapid adoption of open source machine learning (ML) datasets and models exposes today's AI applications to critical risks like data poisoning and supply chain attacks across the ML lifecycle. With growing regulatory pressure to address these issues through greater transparency, ML model vendors face challenges balancing these requirements against confidentiality for data and intellectual property needs. We propose Atlas, a framework that enables fully attestable ML pipelines. Atlas leverages open specifications for data and software supply chain provenance to collect verifiable records of model artifact authenticity and end-to-end lineage metadata. Atlas combines trusted hardware and transparency logs to enhance metadata integrity, preserve data confidentiality, and limit unauthorized access during ML pipeline operations, from training through deployment. Our prototype implementation of Atlas integrates several open-source tools to build an ML lifecycle transparency system, and assess the practicality of Atlas through two case study ML pipelines.

ML安全可追溯性供应链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。