通过地图式分析识别生成模型中的记忆热点,精准干预数据以减少泄露风险。
Data Cartography for Detecting Memorization Hotspots and Guiding Data Interventions in Generative Models
- 基于早期损失和遗忘频率为每个训练样本打分,划分四类区域。
- 仅修剪10%数据,合成密探提取成功率降超40%,性能损失小于0.5点。
- 适合关注模型安全、数据质量与可解释性的研究人员使用。
现代生成模型存在过拟合和意外记忆稀有训练样本的风险,可能被攻击者提取或虚高基准性能。我们提出生成数据制图(GenDataCarto),一种以数据为中心的框架,为每个预训练样本分配难度得分(早期周期损失)和记忆得分(遗忘事件频率),并将其划分为四个象限,指导针对性的数据剪枝与加权调整。我们证明,在平滑性假设下,该记忆得分可下界界定经典影响;且降低高记忆热点权重能通过一致稳定性界显著缩小泛化差距。实验表明,仅修剪10%数据,合成密探提取成功率即下降超过40%,验证集困惑度提升不足0.5。结果表明,基于原理的数据干预可大幅缓解数据泄露,且对生成性能影响极小。
原文摘要 · Abstract (English)
Modern generative models risk overfitting and unintentionally memorizing rare training examples, which can be extracted by adversaries or inflate benchmark performance. We propose Generative Data Cartography (GenDataCarto), a data-centric framework that assigns each pretraining sample a difficulty score (early-epoch loss) and a memorization score (frequency of ``forget events''), then partitions examples into four quadrants to guide targeted pruning and up-/down-weighting. We prove that our memorization score lower-bounds classical influence under smoothness assumptions and that down-weighting high-memorization hotspots provably decreases the generalization gap via uniform stability bounds. Empirically, GenDataCarto reduces synthetic canary extraction success by over 40\% at just 10\% data pruning, while increasing validation perplexity by less than 0.5\%. These results demonstrate that principled data interventions can dramatically mitigate leakage with minimal cost to generative performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。