提出StrTransformer框架,无需编码器即可无监督分离多源信号。
StrTransformer: Source-Wise Structured Transformers for Unsupervised Blind Source Recovery

- 直接优化隐变量矩阵,结合观察空间混合器与源结构分支
- 通过多尺度补丁重建能量实现源级结构正则化,恢复对齐轨迹
- 引入有序多尺度控制器,使各分支专攻不同时间尺度
本文提出StrTransformer,一种面向盲源恢复与分支级隐变量建模的源级结构化Transformer框架。不同于传统编码器推断隐变量,StrTransformer直接联合优化隐源矩阵、观察空间混合器及源级结构Transformer分支。混合器保证重构一致性,每个分支对单一隐源轨迹施加可微结构约束。具体地,将每源转化为多尺度补丁令牌,随机掩码后经局部性偏向Transformer处理,通过掩码补丁重建能量评估,该能量作为隐式源级结构先验。为促使各隐分支专攻不同时间动态,进一步引入有序多尺度控制器,学习分支特异性补丁尺度权重、有序尺度中心与局部注意力斜率。最终目标函数融合观测重构、源级结构正则、模块化辅助惩罚(分离与尺度专一性)。分析表明,该框架具备解耦/耦合结构、受正则化的精确重构纤维,以及由有序分支描述符引起的置换对称性降低。控制实验显示,学习到的分支收敛至不同时间尺度结构,并在事后评估中恢复源对齐的隐轨迹。
原文摘要 · Abstract (English)
This paper proposes StrTransformer, a source-wise structured Transformer framework for blind source recovery and branch-wise latent modeling. Instead of using an encoder to infer latent variables, StrTransformer directly optimizes the latent source matrix together with an observation-space mixer and source-wise structural Transformer branches. The mixer enforces reconstruction consistency, while each Transformer branch imposes a differentiable structural constraint on one latent source trajectory. Specifically, each source is converted into multi-scale patch tokens, randomly masked, processed by a locality-biased Transformer, and evaluated through a masked patch reconstruction energy. This energy acts as an implicit source-wise structural prior. To encourage different latent branches to specialize into different temporal regimes, StrTransformer further introduces an ordered multi-scale controller that learns branch-specific patch-scale weights, ordered scale centers, and locality attention slopes. The resulting objective combines observation reconstruction, source-wise structural regularization, and modular auxiliary penalties for separation and scale specialization. We analyze the decoupling and coupling structure of the objective, the regularized exact-reconstruction fiber, and the reduction of permutation symmetry induced by ordered branch descriptors. A controlled case study shows that the learned branches converge to distinct temporal-scale structures and recover source-aligned latent trajectories under post-hoc evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。