提出一种衡量样本依赖结构匹配度的新指标,可精准检测尾部依赖偏差。
Copula Discrepancy: Benchmarking Dependence Structure
- 基于肯德尔秩相关系数差值设计柯西差异度量,评估样本对目标依赖结构的保持能力。
- 在控制实验中,该指标能有效区分同秩相关但依赖结构不同的样本,且对尾部错配敏感。
- 计算开销仅毫秒级,适合大规模样本的依赖结构诊断,尤其适合生成模型调参场景。
本文研究一种用于评估样本是否准确保留已知二元依赖结构的简单统计量——柯西差异度(Copula Discrepancy, CD)。给定目标柯西族(Clayton 或 Gumbel)及其参数 $θ_P$,CD 比较目标肯德尔秩相关系数 $τ(θ_P)$ 与从样本中拟合出的参数 $ ilde{θ}$ 所对应的 $τ( ilde{θ})$ 之间的绝对差值,即 $|τ(θ_P)−τ( ilde{θ})|$。本文发展了基于矩估计的版本,证明其一致性和渐近正态性,并在独立同分布采样下具有鲁棒性;实证中采用基于最大似然估计的版本以增强对尾部结构误设的检测力。在此基础上,定义了两个信息论型柯西摘要:柯西 KL 散度(CKL)和柯西熵差(CED),并建立了其插值估计量的一致性和中心极限定理。控制实验表明,CD 可可靠区分具有相同 $τ$ 但依赖结构不同的样本,在调节 SGLD 步长时提供依赖感知信号,且在故意制造尾部依赖不匹配时仍保持非零,而朴素 $τ$ 诊断会失效;CKL 与 CED 提供互补的香农式视角,印证上述发现。时间基准测试显示,两种 CD 变体在测试范围内仅增加毫秒级开销,且随样本量近似线性增长,是相比二次代价的核斯坦因差异(KSD)等全局指标更轻量、专注依赖结构的替代方案。
原文摘要 · Abstract (English)
We study a simple statistic for benchmarking how well a sample preserves a known bivariate dependence structure. Given a target copula family (Clayton or Gumbel) and parameter $θ_P$, the Copula Discrepancy (CD) compares the target Kendall's tau $τ(θ_P)$ with the Kendall's tau implied by a parameter $\hatθ$ fitted to the sample within the target family, i.e., $|τ(θ_P)-τ(\hatθ)|$. We develop a moment-based version, prove consistency, asymptotic normality, and robustness results under i.i.d.\ sampling, and use an MLE-based version empirically for greater power against tail-structure misspecification. Building on this, we define two information-theoretic copula summaries, a copula KL divergence (CKL) and a copula entropy gap (CED), and establish basic consistency and central limit results for their plug-in estimators. In controlled experiments, CD reliably separates on-target and off-target copulas with matched Kendall's $τ$, provides a dependence-aware signal for tuning SGLD step sizes where Effective Sample Size favors overly aggressive (and biased) settings, and remains stably nonzero under deliberate tail-dependence mismatch where a naive $τ$-based diagnostic fails; CKL and CED offer a complementary Shannon-style view that echoes these findings. Timing benchmarks show that both CD variants incur only millisecond-level overhead over the tested range and exhibit near-linear empirical scaling in sample size, providing a lightweight, dependence-focused complement to quadratic-cost omnibus discrepancies such as the Kernel Stein Discrepancy (KSD).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。