提出新方法让模型学会用离散符号理解图像,提升推理能力。
Discrete JEPA: Learning Discrete Token Representations without Reconstruction
- 用语义分词和互补目标构建可推理的离散表示
- 在视觉符号预测任务上显著超越现有方法
- 适合研究符号推理与智能体规划的学者
认知智能的核心在于从观测中提取隐藏模式,并利用这些规律系统预测未来。然而,当前图像分词方法在需要符号抽象和逻辑推理的任务中表现有限。为此,我们提出 Discrete-JEPA,通过扩展潜在预测编码框架,引入语义分词和新型互补目标,构建适用于符号推理任务的鲁棒分词表示。Discrete-JEPA 在视觉符号预测任务上显著优于基线模型,且视觉证据显示学习到的语义分词空间中自发涌现出有目的的系统性模式。尽管是初步模型,该方法为推动人工智能系统中的符号世界建模与规划能力具有重要意义。
原文摘要 · Abstract (English)
The cornerstone of cognitive intelligence lies in extracting hidden patterns from observations and leveraging these principles to systematically predict future outcomes. However, current image tokenization methods demonstrate significant limitations in tasks requiring symbolic abstraction and logical reasoning capabilities essential for systematic inference. To address this challenge, we propose Discrete-JEPA, extending the latent predictive coding framework with semantic tokenization and novel complementary objectives to create robust tokenization for symbolic reasoning tasks. Discrete-JEPA dramatically outperforms baselines on visual symbolic prediction tasks, while striking visual evidence reveals the spontaneous emergence of deliberate systematic patterns within the learned semantic token space. Though an initial model, our approach promises a significant impact for advancing Symbolic world modeling and planning capabilities in artificial intelligence systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。