arXiv:2603.15481cs.LGcs.AI2026-03中稿 · 35th International…被引 1

针对表格数据特点,通过增强特征组合覆盖提升无数据知识蒸馏效果。

TabKD: Tabular Knowledge Distillation through Interaction Diversity of Learned Feature Bins

  • 学习与教师决策边界对齐的自适应特征分箱,生成高交互覆盖的合成数据。
  • 在16组配置中14次优于基线,最高达成98.7%学生-教师一致性。
  • 适合隐私敏感场景下的表格模型压缩,尤其适用于缺乏原始数据时。

无数据知识蒸馏可在不使用原始训练数据的情况下实现模型压缩,对隐私敏感的表格领域至关重要。然而,现有方法在表格数据上表现不佳,因其未明确处理特征交互——这是表格模型编码预测知识的根本方式。我们识别出交互多样性(即特征组合的系统性覆盖)是有效表格蒸馏的关键要求。为此,我们提出TabKD:首先学习与教师决策边界对齐的自适应特征分箱,再生成最大化成对交互覆盖的合成查询。在4个基准数据集和4种教师架构上,TabKD在16组配置中的14组实现了最高的学生-教师一致性,显著超越5种先进基线方法。进一步实验表明,交互覆盖程度与蒸馏质量高度相关,验证了核心假设。本工作确立了以交互为导向的探索为表格模型提取的原理性框架。

原文摘要 · Abstract (English)

Data-free knowledge distillation enables model compression without original training data, critical for privacy-sensitive tabular domains. However, existing methods does not perform well on tabular data because they do not explicitly address feature interactions, the fundamental way tabular models encode predictive knowledge. We identify interaction diversity, systematic coverage of feature combinations, as an essential requirement for effective tabular distillation. To operationalize this insight, we propose TabKD, which learns adaptive feature bins aligned with teacher decision boundaries, then generates synthetic queries that maximize pairwise interaction coverage. Across 4 benchmark datasets and 4 teacher architectures, TabKD achieves highest student-teacher agreement in 14 out of 16 configurations, outperforming 5 state-of-the-art baselines. We further show that interaction coverage strongly correlates with distillation quality, validating our core hypothesis. Our work establishes interaction-focused exploration as a principled framework for tabular model extraction.

知识蒸馏表格数据无数据特征交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。