arXiv:2602.01418cs.CVcs.LG2026-02

提出新型视觉位置编码,可推广且适配多种视觉数据

Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General

  • 基于抛物线设计,融合平移/旋转不变性等视觉特性
  • 在ImageNet上比最优基线提升10.5%的外推性能
  • 适用于图像、视频、点云等8个数据集,通用性强

我们提出帕拉博利卡位置编码(PaPE),一种基于抛物线的视觉注意力架构位置编码方法。针对视频、事件相机流、图像或点云等视觉标记序列,旨在编码其位置并充分考虑视觉模态特性。现有工作多将语言模型的一维位置编码扩展至视觉的高维结构,但对视觉特性建模不完整。本文从先验工作中提炼出五大原则:平移不变性、旋转不变性(PaPE-RI)、距离衰减、方向性和上下文感知,构建了具有理论依据的编码方案。在ImageNet-1K上的外推实验表明,PaPE表现显著优于现有方法,绝对性能提升最高达10.5%。跨4类模态的8个数据集的泛化实验显示,PaPE在5个数据集上达到最佳基线性能,在2个数据集上全面超越所有基线,具备良好的通用性。代码已开源。

原文摘要 · Abstract (English)

We propose Parabolic Position Encoding (PaPE), a parabola-based position encoding for vision modalities in attention-based architectures. Given a set of vision tokens-such as from videos, event camera streams, images, or point clouds-our objective is to encode their positions while accounting for the characteristics of vision modalities. Prior works have largely extended position encodings from 1D-sequences in language to nD-structures in vision, but only with partial account of vision characteristics. We address this gap by designing PaPE from principles distilled from prior work: translation invariance, rotation invariance (PaPE-RI), distance decay, directionality, and context awareness. Extrapolation experiments on ImageNet-1K show how PaPE extrapolates remarkably well, improving in absolute terms by up to 10.5\% over the next-best encoding. Generality experiments on 8 datasets across 4 modalities show that PaPE is a general vision position encoding, as PaPE matches the best baseline on 5 datasets and exceeds all on 2 datasets. Code is available at https://github.com/DTU-PAS/parabolic-position-encoding.

位置编码视觉模型注意力机制可推广

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。