arXiv:2509.01181cs.CVcs.AI2025-09AAAI被引 2

动态调整关注区域,实现多主体图像生成的精准控制。

FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus

  • 根据语义复杂度动态识别关键区域,指导生成过程
  • 在多主体生成中显著降低属性泄露,保持主体特征真实
  • 适合需要精细控制的个性化图像生成场景

多主体个性化图像生成旨在无需测试时优化即可合成包含多个指定主体的定制图像。然而,由于难以保持主体保真度并防止跨主体属性泄露,实现对多个主体的细粒度独立控制仍具挑战。本文提出FocusDPO框架,通过动态语义对应关系和参考图像复杂度自适应识别关注区域。训练过程中,方法在噪声时间步上逐步调整焦点区域,采用加权策略奖励信息丰富的区域,惩罚预测置信度低的区域。该框架在DPO过程中依据参考图像的语义复杂度动态调整关注分配,并建立生成图像与参考主体间的鲁棒对应关系。大量实验表明,该方法显著提升现有预训练个性化生成模型的性能,在单主体与多主体个性化图像生成基准上均达到领先水平。方法有效缓解属性泄露问题,同时在多种生成场景下保持优异的主体保真度,推动可控多主体图像合成技术的发展。

原文摘要 · Abstract (English)

Multi-subject personalized image generation aims to synthesize customized images containing multiple specified subjects without requiring test-time optimization. However, achieving fine-grained independent control over multiple subjects remains challenging due to difficulties in preserving subject fidelity and preventing cross-subject attribute leakage. We present FocusDPO, a framework that adaptively identifies focus regions based on dynamic semantic correspondence and supervision image complexity. During training, our method progressively adjusts these focal areas across noise timesteps, implementing a weighted strategy that rewards information-rich patches while penalizing regions with low prediction confidence. The framework dynamically adjusts focus allocation during the DPO process according to the semantic complexity of reference images and establishes robust correspondence mappings between generated and reference subjects. Extensive experiments demonstrate that our method substantially enhances the performance of existing pre-trained personalized generation models, achieving state-of-the-art results on both single-subject and multi-subject personalized image synthesis benchmarks. Our method effectively mitigates attribute leakage while preserving superior subject fidelity across diverse generation scenarios, advancing the frontier of controllable multi-subject image synthesis.

图像生成个性化多主体动态聚焦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。