arXiv:2509.26640cs.LGcs.CR2025-09

通过统计模式分析生成可验证的透明数据卡片,保护隐私同时评估模型鲁棒性。

SPATA: Systematic Pattern Analysis for Detailed and Transparent Data Cards

  • 将表格数据转为无领域依赖的统计模式表示,实现数据脱敏
  • 投影后数据可用于分析特征对模型鲁棒性的影响,结果可复现
  • 适合需数据保密的机构进行模型可信度验证

由于人工智能易受数据扰动和对抗样本影响,部署前必须全面评估模型鲁棒性。然而,通常需访问训练与测试数据集才能分析决策边界和潜在漏洞,可能泄露敏感数据。为提升处理机密数据或关键基础设施组织的透明度,亟需在不披露私有数据的前提下实现外部验证。本文提出系统性模式分析(SPATA),一种确定性方法,将任意表格数据转化为与其领域无关的统计模式表示,生成更详尽、透明的数据卡片。SPATA将每个数据实例投影至离散空间,实现安全分析与比较,避免数据泄露。该投影后的数据可可靠用于评估不同特征对机器学习模型鲁棒性的影响,并生成可解释的行为说明,有助于构建更可信的人工智能。

原文摘要 · Abstract (English)

Due to the susceptibility of Artificial Intelligence (AI) to data perturbations and adversarial examples, it is crucial to perform a thorough robustness evaluation before any Machine Learning (ML) model is deployed. However, examining a model's decision boundaries and identifying potential vulnerabilities typically requires access to the training and testing datasets, which may pose risks to data privacy and confidentiality. To improve transparency in organizations that handle confidential data or manage critical infrastructure, it is essential to allow external verification and validation of AI without the disclosure of private datasets. This paper presents Systematic Pattern Analysis (SPATA), a deterministic method that converts any tabular dataset to a domain-independent representation of its statistical patterns, to provide more detailed and transparent data cards. SPATA computes the projection of each data instance into a discrete space where they can be analyzed and compared, without risking data leakage. These projected datasets can be reliably used for the evaluation of how different features affect ML model robustness and for the generation of interpretable explanations of their behavior, contributing to more trustworthy AI.

数据隐私模型鲁棒性可解释性数据卡片

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。