用积分调制函数实现高效3D位置编码,提升点云建模性能
RelFlexformer: Efficient Attention 3D-Transformers for Integrable Relative Positional Encodings

- 基于非均匀傅里叶变换,设计可灵活集成任意3D相对位置编码的注意力机制
- 计算复杂度降至O(L log L),适用于任意分布的3D点序列
- 特别适合点云建模,实验证明在多个3D数据集上性能更优
我们提出一类新的高效注意力机制,采用由任意可积调制函数 $f$ 定义的通用3D相对位置编码(RPE)。该机制构建了新型3D-Transformer模型,称为 extit{RelFlexformers},可灵活集成这些RPE,且注意力计算时间复杂度为 $O(L /log L)$,其中 $L$ 为输入序列长度。RelFlexformers 基于非均匀傅里叶变换(NU-FFT)理论,自然推广了现有高效RPE-注意力方法,从规则网格中同质嵌入的结构化场景,扩展至一般非结构化异构场景——即令牌在3D空间中任意分布。因此,该模型特别适用于点云建模。我们在大量3D数据集上的广泛实验表明,基于NU-FFT驱动的注意力调制技术显著提升了模型性能。
原文摘要 · Abstract (English)
We present a new class of efficient attention mechanisms applying universal 3D Relative Positional Encoding (RPE) methods given by arbitrary integrable modulation functions $f$. They lead to the new class of 3D-Transformer models, called \textit{RelFlexformers}, flexibly integrating those RPEs, and characterized by the $O(L \log L)$ time complexity of the attention computation for the $L$-length input sequences. RelFlexformers builds on the theory of the Non-Uniform Fourier Transform (NU-FFT), naturally generalizing several existing efficient RPE-attention methods from structured settings with tokens homogeneously embedded in unweighted grids into general non-structured heterogeneous scenarios, where tokens' positions are arbitrarily distributed in the corresponding 3D spaces. As such, RelFlexformers can be applied in particular to model point clouds. Our extensive empirical evaluation on a large portfolio of 3D datasets confirms quality improvements provided by the NU-FFT-driven attention modulation techniques in the RelFlexformers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。