arXiv:2606.31208cs.LGcs.CR2026-06中稿 · ICML被引 1

发现表格大模型在特定条件下会记住训练数据,可能泄露敏感信息。

Probing Memorization of Tabular In-Context Learning

  • 设计新方法分离上下文预测与参数记忆,精准探测记忆行为。
  • 8/10任务中检测到中等记忆信号,最高AUC达0.67,低基数任务最明显。
  • 提醒数据隐私风险,尤其在固定样本微调场景下需加强防护。

大型表格模型(LTMs)通过上下文学习(ICL)在表格任务上达到顶尖性能。尽管大语言模型存在无意记忆训练数据的问题,但表格ICL的记忆机制仍不明确。本文提出ICLMEM探针框架,通过零信息多选题上下文剥离有效上下文模式,迫使模型依赖参数记忆。在受控微调设置下建立成员身份真实标签,控制分布偏移、特征污染、基率谬误等问题,并以预训练模型为参考校准样本难度。对一个主流真实训练的LTM进行评估,在10个任务中有8个检测到中等记忆信号(最高AUC达0.67,1%假阳性率下真阳性率>0.1)。记忆信号在低基数和二分类任务中最强,但在真实训练条件下基本消失。结果表明,仅在单任务微调、固定样本多轮训练且查询量小的情况下出现显著记忆。为保护敏感数据,需采取相应措施。

原文摘要 · Abstract (English)

Large tabular models (LTMs), i.e., tabular foundation models leveraging in-context learning (ICL), achieve state-of-the-art performance on tabular tasks. While LLMs are known to unintentionally memorize training data, the memorization dynamics of LTMs remain largely unexplored. We investigate the potential for parametric memorization in tabular ICL. We introduce ICLMEM, a probing framework designed to separate context-based predictions from parametric memorization. Our zero-information multiple-choice context strips away valid contextual patterns to force the model to fall back on its parametric memory. Our controlled fine-tuning setup establishes membership ground truth and accounts for common pitfalls, e.g., distribution shift, feature contamination, base-rate fallacy, and the pre-trained base model acts as reference to calibrate for sample difficulty. Our controlled evaluation on a leading real-world-trained LTM detects moderate memorization signals in 8 out of 10 tasks ($\text{AUC}$ up to $0.67$ and TPR at $1\%$ FPR $>0.1$). Notably, memorization signals are strongest for low-cardinality and binary tasks. However, they largely vanish under realistic training conditions. Our findings show LTM memorization signals under specific circumstances (single-task fine-tuning with fixed samples across many epochs and small query size). To protect sensitive data, appropriate measures must be taken, which we discuss.

表格模型记忆探测数据隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。