用小波变换统一处理音视频图像,实现跨模态共享令牌结构。
Wavelet as Tokenizer: Preliminary Results on a Shared Wavelet Token Schema for Natural Signals

- 采用一级哈尔小波变换构建共享系数令牌布局
- 跨模态测试中音频最高达39.92 dB,视频23.93 dB PSNR
- 适合多模态信号压缩与低参数量场景的高效建模
本文探讨音频、图像和视频是否可共享一个统一的小波令牌架构,而非依赖各自模态的特定潜在网格。提出一种初步的连续令牌模型,包含一级哈尔小波变换(Haar DWT/IDWT)前端、共享系数令牌布局、可选结构元数据、轻量级模态值适配器及共享令牌编码器-解码器主干。在Speech Commands、EuroSAT RGB和DAVIS 2017数据集上,密集共享模型分别达到39.92 dB(音频)、29.37 dB(图像)和23.93 dB(视频)的PSNR。在连续潜在标量预算下的匹配率扫描显示,视觉性能提升不能仅由潜在容量解释,且附加元数据嵌入并非普适增益源。固定率能量选择提供强非参数基线:在压缩保留比例下,energy_global相比均匀选择平均提升16.73 dB(音频)、16.90 dB(图像)、15.86 dB(视频)。掩码稀疏训练在仅使用50%密集令牌时达到34.45 dB视频PSNR。结果支持统一小波令牌架构与稀疏令牌接口,但尚未建立通用离散词汇表。
原文摘要 · Abstract (English)
This paper studies whether audio, images, and video can share a common wavelet token schema rather than relying on separate modality-specific latent grids. It introduces a preliminary continuous-token model built around a one-level Haar DWT/IDWT frontend, a shared coefficient-token layout, optional structural metadata, lightweight modality value adapters, and a shared token-wise encoder-decoder trunk. On Speech Commands, EuroSAT RGB, and DAVIS 2017 data, a dense shared model reaches 39.92 dB audio, 29.37 dB image, and 23.93 dB video PSNR. A matched-rate sweep under continuous latent scalar budgets indicates that the visual gains are not explained solely by latent capacity, while also showing that additive metadata embeddings are not a universal source of improvement. Finally, fixed-rate energy selection provides a strong non-parametric baseline: energy_global improves average PSNR over uniform selection by 16.73 dB for audio, 16.90 dB for images, and 15.86 dB for video under compressed keep ratios. Masked sparse training reaches 34.45 dB video PSNR with 50% of dense tokens. The results support a unified wavelet token schema and sparse token interface, while stopping short of establishing a universal discrete vocabulary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。