arXiv:2507.17979cs.LG2025-07被引 1

SIFOTL通过统计摘要识别表格数据中的关键变化段,兼顾隐私与噪声鲁棒性。

SIFOTL: A Principled, Statistically-Informed Fidelity-Optimization Method for Tabular Learning

  • 基于隐私安全的统计摘要,用双XGBoost+LLM分离干预信号与噪声
  • 在医保补贴数据上F1达0.85,远超基线方法(最高0.46)
  • 适合医疗数据分析、需解释性的高敏感数据场景

识别表格数据中驱动数据漂移的因素是医疗分析与决策支持系统的重要挑战。隐私规则限制数据访问,复杂过程产生的噪声也阻碍分析。为此,我们提出SIFOTL(Statistically-Informed Fidelity-Optimization Method for Tabular Learning),该方法(i)提取符合隐私要求的数据摘要统计量,(ii)利用双XGBoost模型结合大语言模型辅助,分离干预信号与噪声,(iii)通过帕累托加权决策树融合输出,识别导致漂移的可解释数据子段。相较于现有方法可能忽略噪声或依赖完整数据进行基于大模型的分析,SIFOTL仅使用隐私安全的摘要统计量即可应对双重挑战。在模拟新医保药品补贴的MEPS面板数据集上,SIFOTL取得0.85的F1分数,显著优于BigQuery贡献分析(F1=0.46)和统计检验(F1=0.20)。在基于Synthea ABM生成的18个多样化电子健康记录(EHR)数据集上,无噪声时F1为0.86–0.96,有注入观测噪声时仍保持≥0.75,而基线平均F1仅为0.19–0.67。因此,SIFOTL提供了一种可解释、隐私友好的稳健分析流程。

原文摘要 · Abstract (English)

Identifying the factors driving data shifts in tabular datasets is a significant challenge for analysis and decision support systems, especially those focusing on healthcare. Privacy rules restrict data access, and noise from complex processes hinders analysis. To address this challenge, we propose SIFOTL (Statistically-Informed Fidelity-Optimization Method for Tabular Learning) that (i) extracts privacy-compliant data summary statistics, (ii) employs twin XGBoost models to disentangle intervention signals from noise with assistance from LLMs, and (iii) merges XGBoost outputs via a Pareto-weighted decision tree to identify interpretable segments responsible for the shift. Unlike existing analyses which may ignore noise or require full data access for LLM-based analysis, SIFOTL addresses both challenges using only privacy-safe summary statistics. Demonstrating its real-world efficacy, for a MEPS panel dataset mimicking a new Medicare drug subsidy, SIFOTL achieves an F1 score of 0.85, substantially outperforming BigQuery Contribution Analysis (F1=0.46) and statistical tests (F1=0.20) in identifying the segment receiving the subsidy. Furthermore, across 18 diverse EHR datasets generated based on Synthea ABM, SIFOTL sustains F1 scores of 0.86-0.96 without noise and >= 0.75 even with injected observational noise, whereas baseline average F1 scores range from 0.19-0.67 under the same tests. SIFOTL, therefore, provides an interpretable, privacy-conscious workflow that is empirically robust to observational noise.

表格学习隐私保护可解释性医疗数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。