用算子理论揭示表征可观察性,统一解释模型可解释性与知识迁移的本质。
Platonic Projection Structures: Operator-Induced Observability in Representation Learning

- 通过自伴半正定算子建模观测过程,定义隐状态的不可区分等价类。
- 发现输出式可解释性存在内在局限:核空间成分无法被观测捕捉。
- 适用于理解模型可解释性、知识蒸馏及表征转移中的几何保持机制。
我们通过柏拉图投影结构(PPS)表征表示学习中的可观测性,这是一种基于算子理论的框架,用于分析在部分观测下的表示可达性。系统由三元组 $(H, Π, O)$ 描述,其中 $H$ 是隐表示空间,$Π \succeq 0$ 是观测算子,$O(v)=\langle v,Πv\rangle$ 定义诱导标量可观测量。可观测性由商几何 $H/\ker(Π)$ 表示,刻画在观测下无法区分的隐状态等价类。我们证明量子测量与线性观测模型下的表示推断共享此算子结构,仅算子代数性质不同;对应关系为结构性而非物理性。表征迁移与知识蒸馏亦可解释为近似保持可观测几何,即 $ΦΠ_T \approx Π_S Φ$。PPS还揭示输出式可解释性的结构性限制:位于 $\ker(Π)$ 中的隐成分无法从诱导可观测量中获取,导致归因与解释方法存在内在约束。受控实证验证了核不变可观测性、投影诱导归因差距以及隐空间中秩控制的可观测几何。因此,PPS通过算子诱导的商几何提供了可观测性的显式刻画,并为表示可达性、可解释性与投影推理提供统一视角。
原文摘要 · Abstract (English)
We characterize observability in representation learning through Platonic Projection Structures (PPS), an operator-theoretic framework for analyzing representation accessibility under partial observation. Rather than treating observable outputs as direct reflections of latent representations, PPS models observation through a self-adjoint positive semidefinite operator acting on a latent representation space. A system is represented as a triple $(H, Π, O)$, where $H$ is a latent representation space, $Π\succeq 0$ is an observation operator, and $O(v)=\langle v,Πv\rangle$ defines an induced scalar observable. Observability is characterized by the quotient geometry $H/\ker(Π)$, representing equivalence classes of latent states indistinguishable under observation. We show that quantum measurement and representation inference under linear observation models share this operator-theoretic structure while differing in the algebraic properties of their observation operators; the correspondence is structural rather than physical. Representation transfer and knowledge distillation can likewise be interpreted as approximate preservation of observable geometry through $ΦΠ_T \approx Π_S Φ$. PPS also reveals a structural limitation of output-based interpretability: latent components in $\ker(Π)$ are inaccessible from induced observables, imposing intrinsic constraints on attribution and explanation methods. Controlled empirical validations demonstrate kernel-invariant observability, projection-induced attribution gaps, and rank-controlled observable geometry in latent representation spaces. PPS thus provides an explicit characterization of observability through operator-induced quotient geometry and a unified perspective on representation accessibility, interpretability, and projection-mediated inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。