arXiv:2507.21712stat.MLcs.LG2025-07

从有限样本出发,发现数据点等概率划分实数轴。

An Equal-Probability Partition of the Sample Space: A Non-parametric Inference from Finite Samples

  • 用排序样本点将实线划分为N+1段,每段期望概率均为1/(N+1)
  • 该划分产生离散熵log₂(N+1)比特,反映样本带来的信息量
  • 适用于无需分布假设的密度与尾部估计,提升鲁棒性

本文研究从连续概率分布中抽取的有限样本(大小为N)能揭示什么信息。核心发现是:将N个有序样本点在实线上划分出N+1个区间,每个区间的期望概率质量恰好为1/(N+1)。这一非参数结论源于顺序统计量的基本性质,对任意分布形状均成立。该等概率划分对应的离散熵为log₂(N+1)比特,表征了样本提供的信息量,与香农对连续变量的结果形成对比。本文将此划分框架与传统经验累积分布函数(ECDF)进行比较,探讨其在密度估计和尾部估计中的非参数推断意义。

原文摘要 · Abstract (English)

This paper investigates what can be inferred about an arbitrary continuous probability distribution from a finite sample of $N$ observations drawn from it. The central finding is that the $N$ sorted sample points partition the real line into $N+1$ segments, each carrying an expected probability mass of exactly $1/(N+1)$. This non-parametric result, which follows from fundamental properties of order statistics, holds regardless of the underlying distribution's shape. This equal-probability partition yields a discrete entropy of $\log_2(N+1)$ bits, which quantifies the information gained from the sample and contrasts with Shannon's results for continuous variables. I compare this partition-based framework to the conventional ECDF and discuss its implications for robust non-parametric inference, particularly in density and tail estimation.

非参数推断顺序统计信息熵密度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。