用正弦基重构光谱图像特征,提升对周期性模式的敏感度。
SinBasis Networks: Matrix-Equivalent Feature Extraction for Wave-Like Optical Spectrograms
- 将卷积与注意力重定义为矩阵变换,滤波器权重作为基向量提取特征。
- 在8万张超快光谱图等数据上,重建精度与平移鲁棒性显著提升。
- 适合需要物理先验的波形图像分析,如光谱、音频与周期视频。
波状图像(如阿秒打孔光谱图、光学光谱、音频梅尔频谱图和周期性视频帧)包含关键的谐波结构,传统特征提取器难以捕捉。本文提出统一的矩阵等价框架,将卷积与注意力视为对展平输入的线性变换,揭示滤波器权重作为基向量构成潜在特征子空间。通过在每个权重矩阵上应用逐元素正弦映射,融入频谱先验知识。将此变换嵌入CNN、ViT和胶囊网络,构建出对周期性模式更敏感且具备空间平移不变性的Sin-Basis Networks。在涵盖80,000张合成阿秒打孔光谱图、数千个拉曼、光致发光和FTIR光谱、AudioSet中的梅尔频谱图以及Kinetics中的周期帧的多样化波状图像数据集上,实验显示其在重建精度、平移鲁棒性和零样本跨域迁移方面均有显著提升。通过矩阵同构与Mercer核截断的理论分析,量化证明正弦参数化在数据稀缺情况下增强表达力的同时保持稳定性。Sin-Basis Networks为所有波形成像模态提供了一种轻量级、物理启发的深度学习方法。
原文摘要 · Abstract (English)
Wave-like images--from attosecond streaking spectrograms to optical spectra, audio mel-spectrograms and periodic video frames--encode critical harmonic structures that elude conventional feature extractors. We propose a unified, matrix-equivalent framework that reinterprets convolution and attention as linear transforms on flattened inputs, revealing filter weights as basis vectors spanning latent feature subspaces. To infuse spectral priors we apply elementwise \(\sin(\cdot)\) mappings to each weight matrix. Embedding these transforms into CNN, ViT and Capsule architectures yields Sin-Basis Networks with heightened sensitivity to periodic motifs and built-in invariance to spatial shifts. Experiments on a diverse collection of wave-like image datasets--including 80,000 synthetic attosecond streaking spectrograms, thousands of Raman, photoluminescence and FTIR spectra, mel-spectrograms from AudioSet and cycle-pattern frames from Kinetics--demonstrate substantial gains in reconstruction accuracy, translational robustness and zero-shot cross-domain transfer. Theoretical analysis via matrix isomorphism and Mercer-kernel truncation quantifies how sinusoidal reparametrization enriches expressivity while preserving stability in data-scarce regimes. Sin-Basis Networks thus offer a lightweight, physics-informed approach to deep learning across all wave-form imaging modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。