用自监督学习把视觉信息转成可解释的符号序列。
Extracting Symbolic Sequences from Visual Representations via Self-Supervised Learning
- 扩展DINO框架,用交叉注意力生成符号序列。
- 生成的符号序列能捕捉有意义的抽象层次。
- 适合关注可解释性与高层场景理解的研究者。
本文探索了利用自监督学习(SSL)将复杂的视觉信息抽象为离散、结构化的符号序列的潜力。受语言组织信息以促进推理和泛化方式的启发,我们提出了一种从视觉数据生成符号表示的新方法。为学习这些序列,我们扩展了DINO框架以处理视觉与符号信息。初步实验表明,生成的符号序列具有有意义的抽象层次,尽管仍需进一步优化。该方法的优势在于可解释性:序列由解码器Transformer通过交叉注意力生成,注意力图可关联到特定符号,揭示符号与图像区域的对应关系。该方法为构建可解释的符号表示奠定了基础,具有在高层场景理解中的潜在应用价值。
原文摘要 · Abstract (English)
This paper explores the potential of abstracting complex visual information into discrete, structured symbolic sequences using self-supervised learning (SSL). Inspired by how language abstracts and organizes information to enable better reasoning and generalization, we propose a novel approach for generating symbolic representations from visual data. To learn these sequences, we extend the DINO framework to handle visual and symbolic information. Initial experiments suggest that the generated symbolic sequences capture a meaningful level of abstraction, though further refinement is required. An advantage of our method is its interpretability: the sequences are produced by a decoder transformer using cross-attention, allowing attention maps to be linked to specific symbols and offering insight into how these representations correspond to image regions. This approach lays the foundation for creating interpretable symbolic representations with potential applications in high-level scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。