提出新方法提升噪声数据下的数据蒸馏效果,不依赖干净数据。
Robust Trajectory Distillation: Hybrid Reweighting Meets Teacher-Inspired Targets

- 通过教师模型轨迹分析,动态重加权样本以筛选干净信号。
- 在真实噪声场景下,相比最优基线准确率提升3.2%以上。
- 无需重标注或干净数据,适合工业级小规模训练场景。
数据蒸馏(DD)将大规模数据集压缩为紧凑、信息丰富的子集,以实现高效训练与复用。然而,在存在噪声标签的情况下,传统方法可能将错误关联与有效信号一同压缩,导致鲁棒性下降。现有噪声处理方法(如样本选择、损失加权、标签修正)通常将噪声估计与模型优化强耦合,常需干净样本作为锚点,且易加剧确认偏差——这与DD追求紧凑、即插即用监督的目标相悖。为此,本文提出一种基于轨迹的蒸馏框架,无需重标注或干净子集即可同时抑制噪声并保留可迁移知识。该框架包含两个互补模块:选择性引导重加权(SGR),融合全局遗忘模式(二次分割遗忘)与局部邻域一致性,构建渐进式重加权机制,优先保留教师模型轨迹中的清洁监督信号;教师启发辅助目标(TIAT),从教师模型中间动态中提取辅助残差引导,强化有意义信号的同时保持内部一致性。联合使用SGR与TIAT,可在噪声环境下生成更清洁、更丰富的表示。该方法具有鲁棒性强、标签保真、计算轻量、适用广泛等优点,在对称、非对称及真实世界噪声设置下均显著优于当前最优基线。
原文摘要 · Abstract (English)
Dataset distillation (DD) condenses large corpora into compact, information-rich subsets for efficient training and reuse. However, under noisy supervision, DD risks condensing corrupted associations together with useful signals, degrading robustness. Conventional noisy-label remedies (sample selection, loss weighting, label correction) tightly couple noise estimation with model optimization, often require clean anchors, and can amplify confirmation bias-assumptions that are misaligned with DD's goal of compact, plug-and-play supervision. We therefore propose a trajectory-based DD framework that jointly suppresses noise and preserves transferable knowledge without relabeling or clean subsets. It comprises two complementary components: Selective Guidance Reweighting (SGR), which fuses global forgetting patterns (second-split forgetting) with local neighborhood consistency into a progressive reweighting scheme that prioritizes clean supervision along the teacher trajectory; and Teacher-Inspired Auxiliary Targets (TIAT), which inject auxiliary residual guidance distilled from intermediate teacher dynamics to reinforce informative signals while remaining internally consistent. Together, SGR and TIAT produce distilled datasets with cleaner and richer representations under noisy supervision. The framework is robust, label-preserving, computationally lightweight, and broadly applicable, yielding consistent gains over state-of-the-art DD baselines across symmetric, asymmetric, and real-world noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。