arXiv:2509.05659cs.CV2025-09被引 9

用少量数据实现复杂叙事中角色图像的精准编辑与身份一致

EditIDv2: Editable ID Customization with Data-Lubricated ID Feature Integration for Text-to-Image Generation

  • 通过分解PerceiverAttention,实现低数据依赖的编辑能力注入
  • 在多层语义场景下保持角色身份一致性,支持长文本生成
  • 适合需要高保真角色定制的影视、游戏等复杂场景应用

我们提出EditIDv2,一种无需微调的解决方案,专为高复杂度叙事场景和长文本输入设计。现有角色编辑方法在简单提示下表现良好,但在包含多重语义层、时间逻辑和复杂上下文关系的长文本叙述中,常出现编辑能力下降、语义理解偏差和身份一致性崩溃问题。在EditID的基础上,EditIDv2进一步探索并解决身份特征集成模块的影响。核心在于最小数据润滑下的可编辑性注入。通过精细分解PerceiverAttention,引入身份损失并联合扩散模型进行动态训练,以及离线融合集成模块策略,仅用少量数据即实现复杂叙事环境中的深度多层次语义编辑,同时保持身份一致性。该方法满足长提示与高质量图像生成需求,在IBench评估中取得优异表现。

原文摘要 · Abstract (English)

We propose EditIDv2, a tuning-free solution specifically designed for high-complexity narrative scenes and long text inputs. Existing character editing methods perform well under simple prompts, but often suffer from degraded editing capabilities, semantic understanding biases, and identity consistency breakdowns when faced with long text narratives containing multiple semantic layers, temporal logic, and complex contextual relationships. In EditID, we analyzed the impact of the ID integration module on editability. In EditIDv2, we further explore and address the influence of the ID feature integration module. The core of EditIDv2 is to discuss the issue of editability injection under minimal data lubrication. Through a sophisticated decomposition of PerceiverAttention, the introduction of ID loss and joint dynamic training with the diffusion model, as well as an offline fusion strategy for the integration module, we achieve deep, multi-level semantic editing while maintaining identity consistency in complex narrative environments using only a small amount of data lubrication. This meets the demands of long prompts and high-quality image generation, and achieves excellent results in the IBench evaluation.

图像生成角色编辑长文本理解身份一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。