arXiv:2510.17313cs.LG2025-10NeurIPS被引 3

首个面向多因素时序数据的解耦表征评估基准,推动真实场景下表示学习发展。

Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations

  • 构建六类跨视频、音频与时间序列的标准化多因素时序解耦基准
  • 提出后处理潜空间对齐与柯普曼启发模型,实现当前最优解耦效果
  • 利用视觉语言模型自动标注与零样本评估,减少人工干预

在序列数据中学习解耦表示是深度学习的关键目标,广泛应用于视觉、音频和时间序列领域。尽管真实世界数据包含多个随时间交互的语义因子,但以往研究主要聚焦于简单的两因子静态与动态设置,因数据采集更易而忽视了真实数据的多因子本质。本文首次提出涵盖六个不同数据集(视频、音频、时间序列)的多因素时序解耦标准化评估基准,包含模块化工具链,支持数据集成、模型开发与多因子分析的评估指标。我们还提出一种后处理潜空间探索阶段,可自动将潜空间维度对齐至语义因子,并引入受柯普曼算子启发的模型,达到当前最优性能。此外,我们证明视觉语言模型可自动完成数据标注并作为零样本解耦评估器,无需人工标签与干预。这些贡献共同构建了一个鲁棒且可扩展的多因素时序解耦研究基础。代码已开源于GitHub,数据集与训练模型可在Hugging Face获取。

原文摘要 · Abstract (English)

Learning disentangled representations in sequential data is a key goal in deep learning, with broad applications in vision, audio, and time series. While real-world data involves multiple interacting semantic factors over time, prior work has mostly focused on simpler two-factor static and dynamic settings, primarily because such settings make data collection easier, thereby overlooking the inherently multi-factor nature of real-world data. We introduce the first standardized benchmark for evaluating multi-factor sequential disentanglement across six diverse datasets spanning video, audio, and time series. Our benchmark includes modular tools for dataset integration, model development, and evaluation metrics tailored to multi-factor analysis. We additionally propose a post-hoc Latent Exploration Stage to automatically align latent dimensions with semantic factors, and introduce a Koopman-inspired model that achieves state-of-the-art results. Moreover, we show that Vision-Language Models can automate dataset annotation and serve as zero-shot disentanglement evaluators, removing the need for manual labels and human intervention. Together, these contributions provide a robust and scalable foundation for advancing multi-factor sequential disentanglement. Our code is available on GitHub, and the datasets and trained models are available on Hugging Face.

时序解耦多因素表征基准测试零样本评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。