用代数签名识别概率张量中的结构,无需参数估计。
Algebraic Signatures for Structural Learning in Probability Tensors

- 通过消失二项式作为代数签名,匹配模型结构。
- 在合成与真实语言数据中成功识别出可解释的词组结构。
- 适合对统计模型结构感兴趣的计算语言学研究者。
代数统计通过多项式约束刻画统计模型,但以往多用于解析定义的模型类。本文研究逆问题:从经验概率张量中观察到的消失二项式推断概率结构。将环面模型的消失二项式视为其代数签名,利用代数统计的理想-流形对应关系,构建无需参数估计的结构学习操作流程。通过限制配置矩阵为可计算的克罗内克堆叠类(Kronecker-stack class),使这些签名可显式枚举。在此类中定义最小不变约束(MIC)作为表征每种签名的基本单元,推广了独立性的概念。在合成数据和语料规模的真实语言数据上测试该方法,结果表明所识别的秩一结构对应于可解释的词集。该方法为代数统计在计算语言学中的应用开辟新路径。
原文摘要 · Abstract (English)
Algebraic statistics characterizes statistical models through polynomial constraints, but it has mainly been used for analytically specified model classes. This paper studies the inverse problem: identifying probabilistic structure from vanishing binomials observed in empirical probability tensors. We treat the vanishing binomials of a toric model as its algebraic signature, and turn the ideal-variety correspondence of algebraic statistics into an operational procedure for structural learning that identifies a model by signature matching without parameter estimation. By restricting attention to a computationally tractable class of configuration matrices, which we call {\it the Kronecker-stack class}, we make these signatures explicitly enumerable. Within this class we define minimum invariant constraint (MIC) as the atomic unit characterizing each signature and generalizing the notion of independence. We tested this approach employing MICs on synthetic data as well as on corpus-scale real language data. The results suggested the utility of the method, revealing that the identified rank-one structures correspond to interpretable sets of words. These results open up a new avenue for applying algebraic statistics to computational linguistics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。