用软张量积构建可分布、灵活的视觉组合表征,提升学习效率与性能。
Fully Distributed, Flexible Compositional Visual Representations via Soft Tensor Products
- 提出软张量积(Soft TPR),在分布式框架中实现灵活的组合结构表示。
- 在多个视觉任务中达到最优解耦效果,显著提升收敛速度与小样本表现。
- 适合追求高效、可解释性视觉表征的研究者,尤其关注深度学习与符号系统融合。
自经典主义与联结主义之争以来,系统性组合符号类实体形成组合表征的能力被认为是人类智能的核心。在联结主义系统中,解耦方法因其能生成显式组合表征而备受关注;然而,其依赖于根本上符号化的拼接式结构,与深度学习的连续分布式基础相冲突。为解决这一矛盾,我们扩展了Smolensky的张量积表征(TPR),引入软张量积(Soft TPR),一种在本质上分布式且灵活的组合结构编码方式,并设计了理论严谨的Soft TPR自编码器架构以学习此类表征。在视觉表征学习领域的全面评估表明,该框架在解耦性能上持续优于传统方法——实现最先进的解耦水平,加速表示学习器的收敛,并在下游任务中展现出更优的样本效率和低样本场景表现。这些发现凸显了分布式、灵活组合表征方法的潜力,可能更好地契合深度学习的核心原则,超越传统符号方法。
原文摘要 · Abstract (English)
Since the inception of the classicalist vs. connectionist debate, it has been argued that the ability to systematically combine symbol-like entities into compositional representations is crucial for human intelligence. In connectionist systems, the field of disentanglement has gained prominence for its ability to produce explicitly compositional representations; however, it relies on a fundamentally symbolic, concatenative representation of compositional structure that clashes with the continuous, distributed foundations of deep learning. To resolve this tension, we extend Smolensky's Tensor Product Representation (TPR) and introduce Soft TPR, a representational form that encodes compositional structure in an inherently distributed, flexible manner, along with Soft TPR Autoencoder, a theoretically-principled architecture designed specifically to learn Soft TPRs. Comprehensive evaluations in the visual representation learning domain demonstrate that the Soft TPR framework consistently outperforms conventional disentanglement alternatives -- achieving state-of-the-art disentanglement, boosting representation learner convergence, and delivering superior sample efficiency and low-sample regime performance in downstream tasks. These findings highlight the promise of a distributed and flexible approach to representing compositional structure by potentially enhancing alignment with the core principles of deep learning over the conventional symbolic approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。