arXiv:2604.07522cs.CV2026-04被引 1

无需训练的2D形状编码方法,可高效表示几何与姿态信息

Training-free Spatially Grounded Geometric Shape Encoding (Technical Report)

论文配图:Training-free Spatially Grounded Geometric Shape Encoding (Technical Report)
图 1 · 摘自论文原文
  • 将2D形状分解为单位圆内几何与姿态场,用正交泽尔尼克基编码
  • 编码具备可逆性、自适应性及高频丰富性,支持多种任务应用
  • 适用于需要空间感知的2D智能研究,尤其适合无监督场景

位置编码已成为深度神经网络在离散点位置上建模的通用方法,在一维序列任务中表现卓越。然而,将该思想拓展至二维空间几何形状时,需设计兼顾几何结构、姿态信息及神经网络兼容性的编码策略。本文提出一种无需训练的通用编码方法XShapeEnc,将任意空间定位的2D几何形状映射为紧凑表征,具备可逆性、自适应性、频率丰富性等五项优良特性。具体而言,2D形状被分解为单位圆内的归一化几何与姿态向量,姿态进一步转换为位于单位圆内的谐波姿态场;利用一组正交泽尔尼克基(Zernike bases)对几何与姿态独立或联合编码,并通过频谱传播操作引入高频成分。我们通过广泛分析与实验验证了XShapeEnc的理论合理性、效率、判别能力及适用性,涵盖多种形状感知任务,并基于自建数据集XShapeCorpus进行评估。我们期待XShapeEnc成为迈向二维空间智能研究的重要基础工具。

原文摘要 · Abstract (English)

Positional encoding has become the de facto standard for grounding deep neural networks on discrete point-wise positions, and it has achieved remarkable success in tasks where the input can be represented as a one-dimensional sequence. However, extending this concept to 2D spatial geometric shapes demands carefully designed encoding strategies that account not only for shape geometry and pose, but also for compatibility with neural network learning. In this work, we address these challenges by introducing a training-free, general-purpose encoding strategy, dubbed XShapeEnc, that encodes an arbitrary spatially grounded 2D geometric shape into a compact representation exhibiting five favorable properties, including invertibility, adaptivity, and frequency richness. Specifically, a 2D spatially grounded geometric shape is decomposed into its normalized geometry within the unit disk and its pose vector, where the pose is further transformed into a harmonic pose field that also lies within the unit disk. A set of orthogonal Zernike bases is constructed to encode shape geometry and pose either independently or jointly, followed by a frequency-propagation operation to introduce high-frequency content into the encoding. We demonstrate the theoretical validity, efficiency, discriminability, and applicability of XShapeEnc via extensive analysis and experiments across a wide range of shape-aware tasks and our self-curated XShapeCorpus. We envision XShapeEnc as a foundational tool for research that goes beyond one-dimensional sequential data toward frontier 2D spatial intelligence.

几何编码无训练2D空间泽尔尼克

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。