arXiv:2602.21133cs.LGstat.ML2026-02

让生成模型的离散表示具备可导航的语义拓扑结构,实现直观的人机交互控制。

SOM-VQ: Topology-Aware Tokenization for Interactive Generative Models

  • 用自组织映射改进向量量化,使编码器在低维网格中保持语义相近的邻近关系。
  • 在人体动作生成中实现可学习的序列控制,支持直接调整生成轨迹的语义路径。
  • 适合需要交互式编辑的场景,如舞蹈编排、康复训练和人机交互应用。

向量量化表征虽能支撑强大离散生成模型,但其令牌空间缺乏语义结构,限制了可解释的人类干预。本文提出SOM-VQ,将向量量化与自组织映射结合,学习具有显式低维拓扑结构的离散码本。不同于标准VQ-VAE,SOM-VQ采用拓扑感知更新机制,保留网格上相邻令牌的语义相似性,实现隐空间的直接几何操作。实验表明,SOM-VQ在评估领域生成更易学习的令牌序列,并提供可导航的码空间几何结构。关键优势在于:拓扑组织支持直观的人机协同控制——用户可通过调整令牌空间中的距离来引导生成,实现语义对齐而无需逐帧约束。聚焦人体运动生成,该任务中运动学结构、时间连续性及交互需求使拓扑控制尤为自然,通过简单的网格采样即可实现参考序列的可控发散与收敛。SOM-VQ为音乐、手势等交互式生成任务提供了通用的可解释离散表示框架。

原文摘要 · Abstract (English)

Vector-quantized representations enable powerful discrete generative models but lack semantic structure in token space, limiting interpretable human control. We introduce SOM-VQ, a tokenization method that combines vector quantization with Self-Organizing Maps to learn discrete codebooks with explicit low-dimensional topology. Unlike standard VQ-VAE, SOM-VQ uses topology-aware updates that preserve neighborhood structure: nearby tokens on a learned grid correspond to semantically similar states, enabling direct geometric manipulation of the latent space. We demonstrate that SOM-VQ produces more learnable token sequences in the evaluated domains while providing an explicit navigable geometry in code space. Critically, the topological organization enables intuitive human-in-the-loop control: users can steer generation by manipulating distances in token space, achieving semantic alignment without frame-level constraints. We focus on human motion generation - a domain where kinematic structure, smooth temporal continuity, and interactive use cases (choreography, rehabilitation, HCI) make topology-aware control especially natural - demonstrating controlled divergence and convergence from reference sequences through simple grid-based sampling. SOM-VQ provides a general framework for interpretable discrete representations applicable to music, gesture, and other interactive generative domains.

生成模型离散表示人机交互拓扑结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。