揭示成分数据中单点污染如何通过对数比传播,影响整体分析结果。
Log-Ratio Propagation on the Simplex: A Theory of Cellwise Contamination for Compositional Data
- 基于乘法扰动构建尺度不变污染模型,揭示单个组分污染引发对数比向量的秩一偏移。
- 在单纯形上,单点污染导致估计器的破裂值降低为欧氏情形的 (D-1)/D 倍。
- 污染源可通过变异矩阵的特定行列膨胀模式精准定位,适合地质化学等成分数据分析。
成分数据必须通过对数比分析:尺度不变性是该领域的基本公理,别无选择。中心对数比以所有部分的几何平均数为除数,因此单一污染组分会同时改变所有中心对数比坐标,使对数比向量发生固定位移,任何坐标选择都无法消除。本文围绕这一现象建立单纯形上的单元污染理论。基于乘法扰动的尺度不变污染模型结合传播定理表明,单个原始组分的污染会引发对数比向量的秩一偏移,方向由对比矩阵决定。该扰动模式无法等价于对数比坐标下的独立单元污染模型,因此标准欧氏单元鲁棒方法在单纯形污染机制下不适用。对于其欧氏单元破裂值由列集中配置所体现的估计器(包括MCD、S-、τ-及坐标式M估计量的位置与散度),其在单纯形上的单元破裂值相较欧氏情形降低至(D-1)/D倍,该降幅紧致且源于nD个原始单元与n(D−1)个ilr单元间的归一化差异。变异矩阵的单元影响函数具有诊断指纹特征:单个组分污染仅导致对应行和列被放大,可精确定位污染源。这些结果构成单纯形上单元鲁棒方法的理论基础;配套论文开发了一种利用传播几何的单元鲁棒PCA估计器,并在模拟与地球化学数据上验证。
原文摘要 · Abstract (English)
Compositional data must be analysed through log-ratios: scale invariance, the defining axiom of the field, leaves no alternative. The centred log-ratio divides by the geometric mean of every part, so a single contaminated component shifts every centred-log-ratio coordinate at once, displacing the log-ratio vector by a fixed amount that no choice of coordinates can reduce. We develop a theory of cellwise contamination on the simplex around this observation. A scale-invariant contamination model built from multiplicative perturbation combines with a propagation theorem showing that corruption of a single raw part induces a rank-one shift of the log-ratio vector, with direction determined by the contrast matrix. The resulting perturbation pattern is not equivalent to any independent cellwise contamination model in log-ratio coordinates -- so standard Euclidean cellwise methods applied to log-ratios are ill-posed under the simplex contamination mechanism. For estimators whose Euclidean cellwise breakdown is witnessed by a column-concentrated configuration -- a class including MCD, $S$-, $τ$-, and coordinate-wise $M$-estimators of location and scatter -- the cellwise breakdown value on the simplex is reduced by the factor $(D-1)/D$ relative to its Euclidean counterpart, a reduction that is tight and arises purely from the normalisation mismatch between $nD$ raw cells and $n(D-1)$ ilr cells. The cellwise influence function for the variation matrix carries a diagnostic fingerprint: contamination of a single part inflates exactly one row and column, identifying the responsible component. These results form the theoretical foundation for cellwise-robust methods on the simplex; a companion paper develops a cellwise-robust PCA estimator that exploits the propagation geometry and demonstrates it on simulated and geochemical data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。