arXiv:2511.06610cs.LG2025-11

让数据变成只有指定模型能用的独占品,防止泄露竞争优势。

Non-Rival Data as Rival Products: An Encapsulation-Forging Approach for Data Synthesis

  • 先提取原始数据知识到一个'密钥'模型,再生成专为它设计的合成数据。
  • 合成数据仅对指定模型有效,用少量样本就达到原数据性能。
  • 适合需要共享数据又怕被滥用的企业,保护隐私与竞争优势。

数据的非竞争性给企业带来困境:共享能释放价值,却可能削弱竞争优势。现有数据合成方法常加剧此问题,因生成的数据具有对称效用,任何一方都能获取其价值。本文提出封装伪造(EnFo)框架,一种生成具有非对称效用的竞争性合成数据的新方法。EnFo分两步:首先将原始数据中的预测知识封装进一个指定的“密钥”模型;随后通过优化数据使其刻意过拟合该密钥模型,从而生成合成数据集。这一过程将非竞争性数据转化为竞争性产品,确保其价值仅对目标模型可访问,防止未授权使用,维护数据所有者的竞争优势。实验表明,该框架具备优异的样本效率,仅用原数据的一小部分即可达到相近性能,同时提供强隐私保护和防滥用能力。EnFo为企业在不牺牲核心分析优势的前提下实现战略协作提供了可行方案。

原文摘要 · Abstract (English)

The non-rival nature of data creates a dilemma for firms: sharing data unlocks value but risks eroding competitive advantage. Existing data synthesis methods often exacerbate this problem by creating data with symmetric utility, allowing any party to extract its value. This paper introduces the Encapsulation-Forging (EnFo) framework, a novel approach to generate rival synthetic data with asymmetric utility. EnFo operates in two stages: it first encapsulates predictive knowledge from the original data into a designated ``key'' model, and then forges a synthetic dataset by optimizing the data to intentionally overfit this key model. This process transforms non-rival data into a rival product, ensuring its value is accessible only to the intended model, thereby preventing unauthorized use and preserving the data owner's competitive edge. Our framework demonstrates remarkable sample efficiency, matching the original data's performance with a fraction of its size, while providing robust privacy protection and resistance to misuse. EnFo offers a practical solution for firms to collaborate strategically without compromising their core analytical advantage.

数据合成隐私保护竞争性数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。