arXiv:2605.09572cs.CVcs.AI2026-05中稿 · Neurocomputing

用符号标注生成手语动作,通过多尺度策略提升细节精度。

KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation

论文配图:KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation
图 1 · 摘自论文原文
  • 分阶段生成:先构型后细化,提升整体结构与手指细节。
  • 相比基线模型,动态时间规整误差降低,参数量减少超30%。
  • 适合研究手语生成、高效序列建模的开发者和学者。

从符号化手语标注生成手语动画提供了可扩展的无障碍路径。我们提出KANMultiSign,一种将HamNoSys标注转化为二维人体姿态序列的多尺度序列生成框架。该框架包含两项互补贡献:首先,采用粗到细生成策略并结合多尺度监督——模型先由身体-手部-面部骨架引导以确保全局结构一致性,再精细化手部关节运动以增强指节级细节;其次,探索在Transformer主干中引入柯尔莫哥洛夫-阿诺德网络(KAN)模块,利用可学习的一元函数基元,以紧凑参数化方式建模离散音位符号到连续肢体运动的高非线性映射。在涵盖波兰语、德语、希腊语和法语手语的多个公开数据集上的实验显示,相较强基线模型,本方法在动态时间规整基础上的关节误差持续下降,同时参数量显著减少。受控消融实验进一步表明,结合多尺度监督时,基于KAN的变体能大幅降低参数量且保持竞争力,但其本身并非准确率提升的主要驱动因素。研究结果表明,多尺度监督是提升符号条件姿态生成的核心机制,而KAN则提供了一种高效的建模替代方案。代码将公开。

原文摘要 · Abstract (English)

Sign language production from symbolic notation offers a scalable route to accessible sign animation. We present KANMultiSign, a multi-scale sequence generator that translates HamNoSys notation into two-dimensional human pose sequences. Our framework makes two complementary contributions. First, we introduce a coarse-to-fine generation strategy with multi-scale supervision: the model is first guided by an intermediate body--hand--face scaffold to encourage global structural coherence, and then refines fine-grained hand articulation to improve finger-level detail. Second, we investigate integrating Kolmogorov--Arnold Network modules into a Transformer backbone, using learnable univariate function primitives to model the highly non-linear mapping from discrete phonological symbols to continuous body kinematics with a compact parameterization. Experiments on multiple public corpora spanning Polish, German, Greek, and French sign languages show consistent reductions in dynamic time warping based joint error compared with a strong notation-to-pose baseline, while using substantially fewer parameters. Controlled ablations further indicate that KAN-based variants substantially reduce parameter count while maintaining competitive performance when coupled with multi-scale supervision, rather than serving as the main driver of accuracy gains. These findings position multi-scale supervision as the key mechanism for improving notation-conditioned pose generation, with KAN offering a compact alternative for efficient modeling. Our code will be publicly available.

手语生成序列建模KAN姿态预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。