arXiv:2601.04462cs.LG2026-01

让模型自动学会生成结构,跨数据集共享知识。

Meta-probabilistic Modeling

  • 用分层结构同时捕捉多数据集的共性与个性特征。
  • 在物体中心表征和文本序列建模任务中表现优异。
  • 适合需要跨数据集迁移学习的研究者使用。

概率图模型广泛用于发现数据中的潜在结构,但其效果依赖于合适的模型设计。实际中模型设定困难,常需反复试错,原因在于传统概率图模型通常针对单一数据集。本文提出元概率建模(MPM),针对多个相关数据集的场景,旨在学习生成模型本身的结构。MPM采用分层框架,全局组件编码跨数据集的共享模式,局部参数捕捉各数据集特有的潜在结构。为实现可扩展的学习与推断,我们推导出一种受变分自编码器启发的可处理替代目标函数,并提出双层优化算法。该方法支持多种表达能力强的概率模型,与现有架构如Slot Attention存在联系。在物体中心表征学习和序列文本建模实验中,MPM能有效适应数据并恢复有意义的潜在表示。

原文摘要 · Abstract (English)

Probabilistic graphical models (PGMs) are widely used to discover latent structure in data, but their success hinges on selecting an appropriate model design. In practice, model specification is difficult and often requires iterative trial-and-error. This challenge arises because classical PGMs typically operate on individual datasets. In this work, we consider settings involving collections of related datasets and propose meta-probabilistic modeling (MPM) to learn the generative model structure itself. MPM uses a hierarchical formulation in which global components encode shared patterns across datasets, while local parameters capture dataset-specific latent structure. For scalable learning and inference, we derive a tractable VAE-inspired surrogate objective together with a bi-level optimization algorithm. Our methodology supports a broad class of expressive probabilistic models and has connections to existing architectures, such as Slot Attention. Experiments on object-centric representation learning and sequential text modeling demonstrate that MPM effectively adapts generative models to data while recovering meaningful latent representations.

概率建模元学习生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。