arXiv:2605.16603cs.CV2026-05

通过图约束几何学习,实现多模态生成中属性与身份的精准控制。

Controlla: Learning Controllability via Graph-Constrained Latent Geometry

论文配图:Controlla: Learning Controllability via Graph-Constrained Latent Geometry
图 1 · 摘自论文原文
  • 基于图约束最优传输,将语义属性与身份分离并沿图结构演化。
  • 在43,000样本基准上,身份保持率提升18%,跨模态一致性提高22%。
  • 适合需要高可控性与一致性的图像/视频生成研究者使用。

可控多模态生成通常在推理时通过提示、引导或辅助模块实现,但这类方法未显式建模语义属性的演变过程,易导致身份漂移和跨模态不一致。本文提出Controlla,一种模块化的因子分解控制框架,将可控性视为结构化潜在空间几何的属性。Controlla从多模态输入中学习身份与属性因子,并利用图约束最优传输对齐图先验,促使属性沿图一致轨迹演化,同时保持参考身份不变。为评估该设定,构建了AffectHuman-43K——一个面向参考基情感控制的泄漏感知多模态基准,并引入几何感知指标衡量轨迹一致性与潜在空间解耦性。实验表明,在可控性、身份保持与跨模态对齐方面均有显著提升,且在图敏感性、可扩展性与鲁棒性分析中表现优异。

原文摘要 · Abstract (English)

Controllable multimodal generation is commonly formulated as an inference-time conditioning problem using prompts, guidance, or auxiliary modules. While effective, such approaches do not explicitly structure how semantic attributes evolve, which can lead to identity drift and inconsistent cross-modal behavior. We propose Controlla, a modular factorized-control framework that treats controllability as a property of structured latent geometry. Controlla learns identity and attribute factors from multimodal inputs and aligns them with graph priors using graph-constrained optimal transport, encouraging attributes to follow graph-consistent trajectories while preserving reference identity. To evaluate this setting, we construct AffectHuman-43K, a leakage-aware multimodal benchmark for reference-grounded affective control, and introduce geometry-aware metrics for trajectory consistency and latent disentanglement. Experiments show consistent improvements in controllability, identity preservation, and cross-modal alignment, with additional analyses on graph sensitivity, extensibility, and robustness.

可控生成潜空间几何图结构多模态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。