arXiv:2602.06549cs.LG2026-02被引 1

无需标注即可分离数据噪声与有用特征,提升小样本学习效果

Refining the Information Bottleneck via Adversarial Information Separation

  • 通过自监督对抗机制强制分离有用特征与噪声
  • 多层架构循环利用噪声信息,恢复被误判为噪声的特征
  • 在材料设计等小样本场景中表现优于现有方法

在材料科学等数据稀缺领域,实验数据中的任务相关特征常被测量噪声和实验伪影严重混淆。标准正则化方法难以精确区分有效特征与噪声,而现有对抗性适配方法依赖显式分离标签。为此,我们提出对抗信息分离框架(AdverISF),无需显式监督即可将任务相关特征与噪声分离。AdverISF引入自监督对抗机制,强制任务相关特征与噪声表示在统计上独立,并采用多层分离架构,通过特征层级间循环回收噪声信息,恢复被错误丢弃的有效特征,实现更细粒度的特征提取。大量实验表明,AdverISF在数据稀缺场景下超越现有最先进方法;在真实材料设计任务上的评估也显示其具备更优的泛化性能。

原文摘要 · Abstract (English)

Generalizing from limited data is particularly critical for models in domains such as material science, where task-relevant features in experimental datasets are often heavily confounded by measurement noise and experimental artifacts. Standard regularization techniques fail to precisely separate meaningful features from noise, while existing adversarial adaptation methods are limited by their reliance on explicit separation labels. To address this challenge, we propose the Adversarial Information Separation Framework (AdverISF), which isolates task-relevant features from noise without requiring explicit supervision. AdverISF introduces a self-supervised adversarial mechanism to enforce statistical independence between task-relevant features and noise representations. It further employs a multi-layer separation architecture that progressively recycles noise information across feature hierarchies to recover features inadvertently discarded as noise, thereby enabling finer-grained feature extraction. Extensive experiments demonstrate that AdverISF outperforms state-of-the-art methods in data-scarce scenarios. In addition, evaluations on real-world material design tasks show that it achieves superior generalization performance.

信息瓶颈对抗学习小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。