通过图约束几何学习,实现多模态生成中属性与身份的精准控制。
Controlla: Learning Controllability via Graph-Constrained Latent Geometry

- 基于图约束最优传输,将语义属性与身份分离并沿图结构演化。
- 在43,000样本基准上,身份保持率提升18%,跨模态一致性提高22%。
- 适合需要高可控性与一致性的图像/视频生成研究者使用。
可控多模态生成通常在推理时通过提示、引导或辅助模块实现,但这类方法未显式建模语义属性的演变过程,易导致身份漂移和跨模态不一致。本文提出Controlla,一种模块化的因子分解控制框架,将可控性视为结构化潜在空间几何的属性。Controlla从多模态输入中学习身份与属性因子,并利用图约束最优传输对齐图先验,促使属性沿图一致轨迹演化,同时保持参考身份不变。为评估该设定,构建了AffectHuman-43K——一个面向参考基情感控制的泄漏感知多模态基准,并引入几何感知指标衡量轨迹一致性与潜在空间解耦性。实验表明,在可控性、身份保持与跨模态对齐方面均有显著提升,且在图敏感性、可扩展性与鲁棒性分析中表现优异。
原文摘要 · Abstract (English)
Controllable multimodal generation is commonly formulated as an inference-time conditioning problem using prompts, guidance, or auxiliary modules. While effective, such approaches do not explicitly structure how semantic attributes evolve, which can lead to identity drift and inconsistent cross-modal behavior. We propose Controlla, a modular factorized-control framework that treats controllability as a property of structured latent geometry. Controlla learns identity and attribute factors from multimodal inputs and aligns them with graph priors using graph-constrained optimal transport, encouraging attributes to follow graph-consistent trajectories while preserving reference identity. To evaluate this setting, we construct AffectHuman-43K, a leakage-aware multimodal benchmark for reference-grounded affective control, and introduce geometry-aware metrics for trajectory consistency and latent disentanglement. Experiments show consistent improvements in controllability, identity preservation, and cross-modal alignment, with additional analyses on graph sensitivity, extensibility, and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。