区分数据分布变化中因果机制与协变量偏移,用两个判别器精准量化差异。
Separating Covariate Shift from Mechanism Change with Two Discriminators: CJSD, a Conditional Discrepancy with an Exact Covariate-Concept Decomposition

- 用两个分类器估计条件信息差,分解数据差异为协变量与功能两部分。
- 在202组数据对上,该方法准确分离机制与协变量变化,AUC达1.0。
- 适用于生成模型评估、标注规范检测和公平性审计等场景。
已知输入X后,标签Y还能提供多少关于样本来自哪个数据集的信息?这一量值——可通过两个判别器的保留交叉熵之差D_CJS = CE(Z|X) - CE(Z|X,Y)估算——恰好是协变量偏移无法解释的数据集差异部分。本文提出条件詹森-香农差异(CJSD):引入任务指示符Z,链式规则I(Z;X,Y) = I(Z;X) + I(Z;Y|X)将总任务差异精确分解为协变量轴与功能轴,均可由两个普通分类器估计,无需任务特定预测器、生成模型或自助抽样。证明了协变量零性(纯协变量偏移下功能轴恒为零)、漂移质量定律(D_CJS/ln2等于确定性标签的不一致区域质量),以及单边误设控制不等式(每方向估计误差无条件受限于单一判别器的超额风险),并通过可识别性引理实现固定测度度量化。实验在202个数据对(合成、Electricity、Covertype)上的十项指标测试中,仅两种条件信息估计器(CJSD与kNN插值)实现完全分离,且仅CJSD在维度达256时仍稳定,支持置信区间与序列扩展,同时可审计生成数据的条件保真度、检测输入空间监测忽略的标注指南变更,并用于校准的公平性审计。
原文摘要 · Abstract (English)
After the inputs X are known, how much additional information does the label Y carry about which dataset a sample came from? That single quantity -- estimable as the difference of two discriminators' held-out cross-entropies, D_CJS = CE(Z|X) - CE(Z|X,Y) -- is exactly the part of a dataset difference that covariate shift cannot explain. We propose the Conditional Jensen-Shannon Discrepancy (CJSD): with a task indicator Z, the chain rule I(Z;X,Y) = I(Z;X) + I(Z;Y|X) splits total task discrepancy exactly into a covariate axis and a functional axis, both estimable from two ordinary classifiers, with no task-specific predictors, generative models, or bootstrap surrogates. We prove a covariate-null property (the functional axis is exactly zero under pure covariate shift, however severe), a drift-mass law (D_CJS/ln2 equals the mass of the disagreement region for deterministic labels), a one-sided misspecification-control inequality (each direction of estimation error is bounded, unconditionally, by the excess risk of a single discriminator), and a fixed-measure metrization via an identifiability lemma. Empirically, on a ten-measure battery over 202 dataset pairs (synthetic, Electricity, Covertype), only the two conditional-information estimators -- CJSD and a kNN plug-in for the same estimand -- separate concept from covariate shift with AUC 1.0; the case for CJSD is the estimator: under controlled dimensionality scaling the kNN plug-in fails from d=64 while the discriminator route holds to d=256 with a swappable classifier, and it alone yields paired confidence intervals and sequential extensions from the same learned object. The same estimator audits the conditional fidelity of synthetic-data generators that marginal and joint QA metrics pass, detects annotation-guideline changes invisible to input-space monitors, and supports null-calibrated fairness audits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。