arXiv:2506.09101cs.LGcs.AI2025-06ICML

快速精准定位高维数据中的特征偏移,无需重新训练。

Feature Shift Localization Network

  • 基于神经网络学习数据统计特性,自动定位特征偏移。
  • 在未见过的数据集上仍能准确识别偏移特征,无需重训练。
  • 适用于医疗、金融等异构数据场景,提升下游分析可靠性。

在医疗、生物医学、社会经济、金融、调查及多传感器数据等应用中,数据源间的特征偏移普遍存在,其成因包括异构数据源、噪声测量或处理标准不一致。定位偏移特征有助于查明偏移根源并修正数据,防止下游分析性能下降。尽管已有多种分布偏移检测方法,但特征偏移的精确定位仍具挑战,现有方案或精度不足,或难以扩展至大规模高维数据。本文提出特征偏移定位网络(FSL-Net),一种可快速且准确地在大规模高维数据中定位特征偏移的神经网络。该网络在大量数据集上训练,学习数据的统计特性,能够对未见过的数据集和偏移类型进行定位,且无需重新训练。代码与预训练模型已开源:https://github.com/AI-sandbox/FSL-Net。

原文摘要 · Abstract (English)

Feature shifts between data sources are present in many applications involving healthcare, biomedical, socioeconomic, financial, survey, and multi-sensor data, among others, where unharmonized heterogeneous data sources, noisy data measurements, or inconsistent processing and standardization pipelines can lead to erroneous features. Localizing shifted features is important to address the underlying cause of the shift and correct or filter the data to avoid degrading downstream analysis. While many techniques can detect distribution shifts, localizing the features originating them is still challenging, with current solutions being either inaccurate or not scalable to large and high-dimensional datasets. In this work, we introduce the Feature Shift Localization Network (FSL-Net), a neural network that can localize feature shifts in large and high-dimensional datasets in a fast and accurate manner. The network, trained with a large number of datasets, learns to extract the statistical properties of the datasets and can localize feature shifts from previously unseen datasets and shifts without the need for re-training. The code and ready-to-use trained model are available at https://github.com/AI-sandbox/FSL-Net.

特征偏移神经网络数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。