arXiv:2510.18109cs.CRcs.LG2025-10

保护隐私的数据评估协议,让买卖双方不暴露秘密也能测数据价值。

PrivaDE: Privacy-preserving Data Evaluation for Blockchain-based Data Marketplaces

  • 双方联合计算数据效用分,不泄露模型参数和原始数据。
  • 百万参数模型在线评估仅需15分钟内完成。
  • 适合去中心化数据交易市场,保障公平与安全。

在区块链数据市场中,购买数据前评估其价值对训练高质量机器学习模型至关重要,但模型开发者与数据提供方往往不愿暴露各自专有资产。我们提出 PrivaDE,一种隐私保护协议,使模型所有者与数据所有者能在不完全披露模型参数、原始特征或标签的前提下,联合计算候选数据集的效用评分。PrivaDE 具备强安全性,可抵御恶意行为,并能集成至基于区块链的市场,由智能合约确保公平执行与支付。为提升实用性,我们提出了高效的安全模型推理优化方案,以及一种仅需少量代表性数据子集即可反映其下游训练影响的模型无关评分方法。实验表明,即便面对拥有数百万参数的模型,PrivaDE 的在线运行时间仍控制在15分钟以内。本工作为去中心化机器学习生态中的公平、自动化数据市场奠定了基础。

原文摘要 · Abstract (English)

Evaluating the usefulness of data before purchase is essential when obtaining data for high-quality machine learning models, yet both model builders and data providers are often unwilling to reveal their proprietary assets. We present PrivaDE, a privacy-preserving protocol that allows a model owner and a data owner to jointly compute a utility score for a candidate dataset without fully exposing model parameters, raw features, or labels. PrivaDE provides strong security against malicious behavior and can be integrated into blockchain-based marketplaces, where smart contracts enforce fair execution and payment. To make the protocol practical, we propose optimizations to enable efficient secure model inference, and a model-agnostic scoring method that uses only a small, representative subset of the data while still reflecting its impact on downstream training. Evaluation shows that PrivaDE performs data evaluation effectively, achieving online runtimes within 15 minutes even for models with millions of parameters. Our work lays the foundation for fair and automated data marketplaces in decentralized machine learning ecosystems.

数据市场隐私计算区块链机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。