arXiv:2604.02482cs.LG2026-04

提出结构外推数据生成框架,让模型生成训练数据之外的新内容。

SEDGE: Structural Extrapolated Data Generation

论文配图:SEDGE: Structural Extrapolated Data Generation
图 1 · 摘自论文原文
  • 基于数据生成过程的假设,设计结构感知的优化与扩散采样方法。
  • 在保守假设下可近似识别新数据分布,否则无法唯一确定分布。
  • 在合成数据和图像生成中验证了生成效果,适用于需要突破训练范围的场景。

本文针对训练数据之外的数据生成难题,提出基于数据生成过程合理假设的结构外推数据生成框架(SEDGE)。我们给出了在何种条件下可可靠生成符合新规格的数据,并在某些“保守”假设下实现了此类数据分布的近似可识别性,而无这些假设时则存在固有的不可识别性。算法层面,提出了基于结构信息优化策略或扩散后验采样的实用生成方法。通过合成数据实验以及真实场景下的图像外推生成,验证了该框架的有效性。

原文摘要 · Abstract (English)

This paper aims to address the challenge of data generation beyond the training data and proposes a framework for Structural Extrapolated Data GEneration (SEDGE) based on suitable assumptions on the underlying data-generating process. We provide conditions under which data satisfying novel specifications can be generated reliably, together with the approximate identifiability of the distribution of such data under certain ``conservative" assumptions, as well as the inherent non-identifiability of this distribution without such assumptions. On the algorithmic side, we develop practical methods to achieve extrapolated data generation, based on a structure-informed optimization strategy or diffusion posterior sampling, respectively. We verify the extrapolation performance on synthetic data and also consider extrapolated image generation as a real-world scenario to illustrate the validity of the proposed framework.

数据生成结构外推扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。