提出新型位置编码STRING,高效实现2D/3D坐标不变性,提升视觉模型性能。
Learning the RoPEs: Better 2D and 3D Position Encodings with STRING

- 基于旋转编码扩展,设计可分离的平移不变位置编码
- 在机器人视觉任务中显著提升开放词汇目标检测效果
- 理论证明通用性,适合需要3D建模的视觉与控制场景
我们提出STRING:可分离的平移不变位置编码。STRING通过统一的理论框架扩展了旋转位置编码(RoPE),该编码近期被广泛应用于大语言模型中。重要的是,STRING仍保持精确的平移不变性,适用于任意维度的标记坐标,同时维持低计算开销。这一特性在机器人领域尤为重要,因高效3D标记表示是关键。我们将STRING集成到基于RGB(-D)输入(颜色加可选深度)的视觉变换器中,在开放词汇目标检测和机器人控制器任务中均取得显著提升。我们通过严谨的数学分析证明了方法的普遍性。
原文摘要 · Abstract (English)
We introduce STRING: Separable Translationally Invariant Position Encodings. STRING extends Rotary Position Encodings, a recently proposed and widely used algorithm in large language models, via a unifying theoretical framework. Importantly, STRING still provides exact translation invariance, including token coordinates of arbitrary dimensionality, whilst maintaining a low computational footprint. These properties are especially important in robotics, where efficient 3D token representation is key. We integrate STRING into Vision Transformers with RGB(-D) inputs (color plus optional depth), showing substantial gains, e.g. in open-vocabulary object detection and for robotics controllers. We complement our experiments with a rigorous mathematical analysis, proving the universality of our methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。