arXiv:2512.15955cs.LG2025-12

分析决策树论文中使用的预测变量,发现多数涉及受监管数据。

Governance by Evidence: Regulated Predictors in Decision-Tree Models

  • 用论文作为样本,识别决策树中的预测变量所属的受监管类别
  • 医疗数据类预测变量占比最高,不同行业差异显著
  • 为机器学习实践提供隐私合规参考,尤其适合政策制定者

决策树方法广泛应用于结构化表格数据,在多个领域因其可解释性备受青睐。然而,现有研究常披露所用预测变量(如年龄、诊断代码、位置等),而这些数据类型日益受到隐私法规约束。本文以已发表的决策树论文为真实世界使用受监管数据的代理,构建了决策树研究语料库,并将每个报告的预测变量归类至受监管数据类别(如健康数据、生物识别信息、儿童数据、金融属性、位置轨迹、政府身份标识)。随后,将各类别与欧盟和美国隐私法律的具体条文关联。结果显示,许多报告的预测变量属于受监管范畴,其中医疗数据占比最大,且不同行业间存在明显差异。我们分析了各类别在时间上的分布、行业构成及演变趋势,并基于各法律框架的参考年份总结合规时间线。研究证据支持采用隐私保护方法和治理审查机制,可为超越决策树的机器学习实践提供指导。

原文摘要 · Abstract (English)

Decision-tree methods are widely used on structured tabular data and are valued for interpretability across many sectors. However, published studies often list the predictors they use (for example age, diagnosis codes, location). Privacy laws increasingly regulate such data types. We use published decision-tree papers as a proxy for real-world use of legally governed data. We compile a corpus of decision-tree studies and assign each reported predictor to a regulated data category (for example health data, biometric identifiers, children's data, financial attributes, location traces, and government IDs). We then link each category to specific excerpts in European Union and United States privacy laws. We find that many reported predictors fall into regulated categories, with the largest shares in healthcare and clear differences across industries. We analyze prevalence, industry composition, and temporal patterns, and summarize regulation-aligned timing using each framework's reference year. Our evidence supports privacy-preserving methods and governance checks, and can inform ML practice beyond decision trees.

隐私保护决策树数据合规监管

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。