arXiv:2512.23335cs.CVcs.LG2025-12被引 1

视觉理解需将感知变体归约为离散语义状态,模型须具备拓扑重构能力。

Visual Language Hypothesis

  • 从语义语言假设出发,构建视觉空间的纤维丛结构,区分语义与干扰因素。
  • 语义不变性依赖非同胚的判别目标,如标签监督或跨实例对齐。
  • 模型需支持'扩展-收缩'拓扑变换,才能实现有效的语义抽象。

我们从结构与拓扑角度研究视觉表征学习。核心假设是:视觉理解预设了视觉语义语言,即大量感知观测对应少数离散语义状态。结合表征学习中广泛接受的可迁移性与抽象性前提,该假设意味着视觉观察空间应具有纤维丛结构,其中噪声变化分布在纤维上,语义对应于商基空间。由此推导出两个理论结论:其一,商空间 X/G 不是原始空间 X 的子流形,无法仅通过光滑变形获得;语义不变性要求非同胚的判别目标,如标签监督、跨实例识别或跨模态对齐以提供显式语义等价。其二,逼近商空间对模型架构有结构性要求:语义抽象不仅需要外部语义目标,还需支持拓扑变化的表示机制——即先几何展开以分离结构,再收缩形成离散语义区域的‘扩展-收缩’过程。这些结果为解释而非规定性,框架与大规模判别模型和多模态模型的经验规律以及统计学习理论的经典原则一致。

原文摘要 · Abstract (English)

We study visual representation learning from a structural and topological perspective. We begin from a single hypothesis: that visual understanding presupposes a semantic language for vision, in which many perceptual observations correspond to a small number of discrete semantic states. Together with widely assumed premises on transferability and abstraction in representation learning, this hypothesis implies that the visual observation space must be organized in a fiber bundle like structure, where nuisance variation populates fibers and semantics correspond to a quotient base space. From this structure we derive two theoretical consequences. First, the semantic quotient X/G is not a submanifold of X and cannot be obtained through smooth deformation alone, semantic invariance requires a non homeomorphic, discriminative target for example, supervision via labels, cross-instance identification, or multimodal alignment that supplies explicit semantic equivalence. Second, we show that approximating the quotient also places structural demands on the model architecture. Semantic abstraction requires not only an external semantic target, but a representation mechanism capable of supporting topology change: an expand and snap process in which the manifold is first geometrically expanded to separate structure and then collapsed to form discrete semantic regions. We emphasize that these results are interpretive rather than prescriptive: the framework provides a topological lens that aligns with empirical regularities observed in large-scale discriminative and multimodal models, and with classical principles in statistical learning theory.

视觉表征拓扑学习语义抽象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。