arXiv:2608.25138cs.LGcs.AI2026-08

用条件后验流匹配统一生成与表征学习,实现更准确的多模态补全。

Drift Variation Autoencoder: Unifying Generation and Representation Learning through Conditional Posterior Flow Matching

  • 以条件后验分布为统一目标,通过流匹配训练编码器与解码器
  • 在CrossGeom-4上线性探测R²达0.9990~0.9992,误差提升13.5~15.7倍
  • 适合多模态数据生成与表征学习场景,尤其关注完整模态补全

随机掩码、裁剪或模态缺失使确定性重建成为不完整的目标:一个观测可能对应多种清洁补全。本文将条件后验 $P(X ext{∣}C)$ 视为条件生成与生成充分表征学习的共同统计对象。漂移变分自编码器训练一个掩码编码器 $Z=E(C)$ 和一个条件流解码器,使用单一清洁预测流匹配损失。分析首先将理想条件KL分解为生成器近似误差和表征不足项 $I(X;C ext{∣}Z)$。随后推导出条件流匹配的正交风险分解。对于仿射高斯路径,若且仅当 $P(X ext{∣}Z)=P(X ext{∣}C)$ 时,清洁预测表征差距为零。因此,流匹配引起的编码器依赖超额清洁预测风险与理想条件KL的零集一致,尽管数值不等。具有零噪声终点的精确条件场在联合最优下生成 $P(X ext{∣}Z)$,从而生成 $P(X ext{∣}C)$。该结果在连续多模态乘积空间中仍成立,前提是完整模态元组始终作为每种观测掩码的流目标。在CrossGeom-4的18次控制基准测试中,可观测因素的线性探测 $R^2$ 达 $0.9990$-$0.9992$,打乱联合模型编码器条件使条件误差增加 $13.5$-$15.7$ 倍,联合目标注意力使共享未观测因素上的分歧降低 $90.1$-$92.8 ext{%}$,相较独立目标解码器。可见模态也可被生成与重建,直接验证全元组目标。无条件模式平衡仍不完美,限制了实证主张为受控多模态概念验证。

原文摘要 · Abstract (English)

Stochastic masking, cropping, or modality removal makes deterministic reconstruction an incomplete target: one observation can admit many clean completions. This work takes the corresponding posterior $P(X\mid C)$ as the common statistical object for conditional generation and generatively sufficient representation learning. Drift Variation autoencoder trains a masked encoder $Z=E(C)$ and a conditional flow decoder with one clean-prediction Flow Matching loss. The analysis first decomposes the ideal conditional KL into generator approximation and the representation deficiency $I(X;C\mid Z)$. It then derives orthogonal risk decompositions for conditional Flow Matching. For an affine Gaussian path, the clean-prediction representation gap is zero if and only if $P(X\mid Z)=P(X\mid C)$. Thus the encoder-dependent excess clean-prediction risk induced by Flow Matching and the profiled ideal conditional KL have the same posterior-sufficient zero set, without being numerically equal objectives. An exact conditional field with a zero-noise endpoint then generates $P(X\mid Z)$ and hence $P(X\mid C)$ at a joint ideal optimum. The result extends to continuous multimodal product spaces when the complete modality tuple remains the Flow target for every observation mask. On CrossGeom-4, an 18-run controlled benchmark, observable factors have linear-probe $R^2$ of $0.9990$-$0.9992$, shuffling the joint model's encoder condition increases conditional error by $13.5\times$-$15.7\times$, and joint target attention reduces disagreement on an unobserved factor shared by two outputs by $90.1$-$92.8\%$ relative to independent target decoders. Visible modalities are also generated and reconstructed, directly validating the full-tuple objective. Unconditional mode balance remains imperfect, delimiting the empirical claim to a controlled multimodal proof of concept.

生成模型表征学习流匹配多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。