arXiv:2604.17568cs.LGmath.ST2026-04

在无监督条件下,通过集合运算恢复隐变量的结构关系。

Diverse Dictionary Learning

论文配图:Diverse Dictionary Learning
图 1 · 摘自论文原文
  • 利用观察数据的集合运算(交、补、对称差)推断隐变量结构。
  • 即使无强假设,仍可保证部分隐变量关系的可识别性。
  • 该方法适用于真实数据,且可轻松融入现有模型框架。

仅给定观测数据 $X = g(Z)$,其中隐变量 $Z$ 和生成过程 $g$ 均未知时,恢复 $Z$ 是病态问题。现有方法常依赖线性假设或辅助监督,但这些假设在实际中难以验证,理论保证也易被轻微违背破坏。为此,本文提出多样字典学习问题,从互补视角出发:在无法完全识别的通用场景下,仍能可靠恢复什么?我们证明,在无强假设前提下,与任意观测关联的隐变量的交集、补集、对称差,以及隐变量与观测之间的依赖结构,仍可被识别,仅存在适当不确定性。这些集合论结果可通过集合代数组合,构建出如种差定义等结构化、本质化的隐藏世界视图。当存在足够结构多样性时,可进一步实现所有隐变量的完整可识别性。所有可识别性收益均来自估计阶段的一个简单归纳偏置,该偏置可直接集成到多数模型中。我们在合成与真实数据上验证了理论并展示了该偏置的优势。

原文摘要 · Abstract (English)

Given only observational data $X = g(Z)$, where both the latent variables $Z$ and the generating process $g$ are unknown, recovering $Z$ is ill-posed without additional assumptions. Existing methods often assume linearity or rely on auxiliary supervision and functional constraints. However, such assumptions are rarely verifiable in practice, and most theoretical guarantees break down under even mild violations, leaving uncertainty about how to reliably understand the hidden world. To make identifiability actionable in the real-world scenarios, we take a complementary view: in the general settings where full identifiability is unattainable, what can still be recovered with guarantees, and what biases could be universally adopted? We introduce the problem of diverse dictionary learning to formalize this view. Specifically, we show that intersections, complements, and symmetric differences of latent variables linked to arbitrary observations, along with the latent-to-observed dependency structure, are still identifiable up to appropriate indeterminacies even without strong assumptions. These set-theoretic results can be composed using set algebra to construct structured and essential views of the hidden world, such as genus-differentia definitions. When sufficient structural diversity is present, they further imply full identifiability of all latent variables. Notably, all identifiability benefits follow from a simple inductive bias during estimation that can be readily integrated into most models. We validate the theory and demonstrate the benefits of the bias on both synthetic and real-world data.

隐变量建模可识别性集合运算无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。