arXiv:2505.02823cs.CV2025-05被引 11

仅用单人数据实现多人定制,解决数据少与属性混淆问题

MUSAR: Exploring Multi-Subject Customization from Single-Subject Dataset via Attention Routing

  • 通过双图对构建技术突破单人数据限制
  • 动态注意力路由使多人表征解耦且泛化能力强
  • 无需多人训练数据仍超越现有方法性能

当前多主体定制方法面临两大挑战:难以获取多样化的多主体训练数据,以及不同主体间属性纠缠。为此,我们提出MUSAR——一种仅需单主体训练数据即可实现稳健多主体定制的简单而有效框架。首先,为突破数据限制,引入无偏双联学习,从单主体图像构建双联训练对以促进多主体学习,并通过静态注意力路由和双分支LoRA主动校正双联构造带来的分布偏差。其次,为消除跨主体纠缠,引入动态注意力路由机制,自适应建立生成图像与条件主体之间的双射映射。该设计不仅实现多主体表征解耦,还能在参考主体数量增加时保持可扩展的泛化性能。大量实验表明,尽管仅使用单主体数据,MUSAR在图像质量、主体一致性及交互自然度上均优于现有方法(包括在多主体数据上训练的模型)。

原文摘要 · Abstract (English)

Current multi-subject customization approaches encounter two critical challenges: the difficulty in acquiring diverse multi-subject training data, and attribute entanglement across different subjects. To bridge these gaps, we propose MUSAR - a simple yet effective framework to achieve robust multi-subject customization while requiring only single-subject training data. Firstly, to break the data limitation, we introduce debiased diptych learning. It constructs diptych training pairs from single-subject images to facilitate multi-subject learning, while actively correcting the distribution bias introduced by diptych construction via static attention routing and dual-branch LoRA. Secondly, to eliminate cross-subject entanglement, we introduce dynamic attention routing mechanism, which adaptively establishes bijective mappings between generated images and conditional subjects. This design not only achieves decoupling of multi-subject representations but also maintains scalable generalization performance with increasing reference subjects. Comprehensive experiments demonstrate that our MUSAR outperforms existing methods - even those trained on multi-subject dataset - in image quality, subject consistency, and interaction naturalness, despite requiring only single-subject dataset.

多主体定制注意力路由单主体数据表征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。