arXiv:2510.05717cs.LG2025-10被引 2

用扩散模型实现跨模态无监督时序解耦,效果优于现有方法。

DiffSDA: Unsupervised Diffusion Sequential Disentanglement Across Modalities

  • 基于潜变量扩散建模,统一处理多模态时序数据。
  • 在真实世界数据上显著优于当前最佳方法,提升解耦精度。
  • 提供严格评估协议,适合研究多模态时序建模的学者。

无监督表示学习,尤其是时序解耦,旨在不依赖标签的情况下分离数据中的静态与动态变化因素。该问题仍具挑战性,因为现有基于变分自编码器和生成对抗网络的方法通常依赖多个损失项,使优化过程复杂化。此外,时序解耦方法在真实数据上应用时面临困难,且尚无标准化的评估协议用于衡量其性能。最近,扩散模型已成为最先进的生成模型,但尚未有理论框架支持其在时序解耦中的应用。本文提出扩散时序解耦自编码器(DiffSDA),一种适用于多种真实世界数据模态(包括时间序列、视频和音频)的新型、模态无关框架。DiffSDA利用新的概率建模方式、潜变量扩散机制及高效采样器,并引入具有挑战性的评估协议以进行严格测试。在多样化的现实基准上的实验表明,DiffSDA在时序解耦方面显著优于近期最先进的方法。

原文摘要 · Abstract (English)

Unsupervised representation learning, particularly sequential disentanglement, aims to separate static and dynamic factors of variation in data without relying on labels. This remains a challenging problem, as existing approaches based on variational autoencoders and generative adversarial networks often rely on multiple loss terms, complicating the optimization process. Furthermore, sequential disentanglement methods face challenges when applied to real-world data, and there is currently no established evaluation protocol for assessing their performance in such settings. Recently, diffusion models have emerged as state-of-the-art generative models, but no theoretical formalization exists for their application to sequential disentanglement. In this work, we introduce the Diffusion Sequential Disentanglement Autoencoder (DiffSDA), a novel, modal-agnostic framework effective across diverse real-world data modalities, including time series, video, and audio. DiffSDA leverages a new probabilistic modeling, latent diffusion, and efficient samplers, while incorporating a challenging evaluation protocol for rigorous testing. Our experiments on diverse real-world benchmarks demonstrate that DiffSDA outperforms recent state-of-the-art methods in sequential disentanglement.

扩散模型时序解耦无监督学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。