arXiv:2606.07914stat.MLcs.LG2026-06

在无标签混合模型中,利用坐标独立性实现成分与混合矩阵的可识别恢复。

Identifiability and Estimation for Unlabeled Finite Mixtures under Marginal Independence

论文配图:Identifiability and Estimation for Unlabeled Finite Mixtures under Marginal Independence
图 1 · 摘自论文原文
  • 基于坐标对独立性假设,通过仿射组合重构潜在成分。
  • 在特定秩条件下,可完全识别所有成分并估计混合矩阵。
  • 提出PM-MMD估计算法,适用于流式细胞等真实数据场景。

研究无标签有限混合模型中成分恢复与混合矩阵估计问题,其可观测分布共享相同潜成分但混合权重未知。主要识别信号为边际独立性:每个成分至少在一个坐标对上独立,且不提供标签、干净成分样本或混合权重。首先证明乘积成分的结构结果:在单变量边缘跨度的子集秩条件下,任何独立仿射组合必等于单一成分。进而将该原理扩展至可观测混合模型,在对应子集秩、满秩及无抵消条件下,边际独立的仿射组合可恢复对应潜成分。当每个成分均在某坐标对上独立时,所有成分均可识别,混合矩阵在给定完成条件下可恢复。最后提出基于仿射组合的乘积-边际最大均值差异(PM-MMD)估计算法,并证明在近似边际独立下的统一收敛性和稳定性。该框架分离了假设的实证角色:不可约性通常无法仅从无标签混合中检验,而边际独立可通过保留的PM-MMD提供候选级别诊断。控制实验与流式细胞数据验证了边际独立作为有效恢复信号的作用。在多成分比较中,条件感知的代表性选择使PM-MMD更稳定,优于聚类、分解及成对混合比例基线。

原文摘要 · Abstract (English)

We study component recovery and mixing-matrix estimation from unlabeled finite mixtures whose observable distributions share the same latent components but have unknown mixing weights. The main identifying signal is marginal independence: each component is assumed to be independent on at least one coordinate pair, but no labels, clean component samples, or mixing weights are observed. We first prove a structural result for product components: under a subset-rank condition on the spans of the univariate marginals, any independent affine combination of the components must coincide with a single component. We then extend this principle to observable mixtures and show that, under the corresponding subset-rank, full-rank, and no-cancellation conditions, marginally independent affine combinations recover the corresponding latent components. When every component is independent on some coordinate pair, all components are identifiable, and the mixing matrix is recoverable under the stated completion conditions. Finally, we propose a Product-Marginal Maximum Mean Discrepancy (PM-MMD) estimator over affine combinations of the observable mixtures and prove uniform convergence and stability under approximate marginal independence. This framework also separates the empirical roles of the assumptions: irreducibility is, in general, not directly testable from the unlabeled mixtures alone, whereas marginal independence yields a candidate-level diagnostic through held-out PM-MMD. Controlled and flow-cytometry experiments show when marginal independence provides a useful recovery signal. In the reported multi-component comparisons, condition-aware representative selection stabilizes PM-MMD and improves recovery relative to clustering, factorization, and pairwise mixture-proportion baselines using the same unlabeled mixtures.

混合模型成分识别边际独立PM-MMD

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。