研究水印对离散分布估计的影响,发现加水印无法提升性能除非误检率归零。
Minimax bounds for watermarked and masked recursive discrete distribution estimation
- 通过最小最大损失分析水印对递归分布估计的影响
- 证明当真实样本比例趋近零时,加水印无法提升性能
- 提出随机掩码方法缩小性能差距,适合关注数据真实性验证的研究者
水印被提出用于在无元数据情况下识别合成样本,但其确切影响尚不明确。在缺乏区分机制时,已有研究显示添加合成样本会显著降低新真实样本的边际效用。本文研究了存在水印情况下的递归离散分布估计最小最大损失,与无辅助和有预言辅助的情况对比。当真实样本占比渐近趋于零时,我们给出了下界,表明除非检测的漏报率也趋于零,否则无法通过加水印改善性能。此外,在多数情形下,一系列简单确定性估计器的最坏情况损失与对应下界相差常数倍。最后,我们提出掩码这一随机化过程,将剩余情形中的差距缩小至詹森差。我们猜想更紧的下界论证可进一步弥合该差距。
原文摘要 · Abstract (English)
Watermarking has been proposed as a way to identify synthetic samples in estimation settings where no metadata is available to distinguish them from real samples, but its precise effects remain unexplored. In the absence of a distinguishing mechanism, it has been shown that adding synthetic samples significantly reduces the marginal efficacy of new real samples. In this work, we study the minimax loss of such recursive discrete distribution estimation in the presence of watermarks in contrast to the unassisted and oracle-assisted losses. When the fraction of real samples vanishes asymptotically, we provide a lower bound that shows that it is impossible to improve performance by adding watermarks unless the false negative rate of detection also vanishes. Additionally, we show that in most regimes, the worst-case losses of a sequence of simple deterministic estimators match the corresponding lower bounds up to constants. Finally, we propose masking, a randomization procedure that narrows the gap in the remaining regimes to a Jensen gap. We conjecture that a tighter lower bound argument can close this gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。