arXiv:2511.11989cs.CV2025-11

突破人脸特写限制,实现身份一致的个性化场景生成

BeyondFacial: Identity-Preserving Personalized Generation Beyond Facial Close-ups

  • 采用双路推理架构分离身份与语义特征
  • 在噪声预测阶段动态融合身份信息,避免干扰语义表达
  • 无需手动标记或微调,适合影视级角色场景创作

身份保持的个性化生成(IPPG)已推动影视制作与艺术创作发展,但现有方法过度依赖人脸区域,导致输出多为面部特写。这些方法在复杂文本提示下存在视觉叙事性弱、语义一致性差的问题,核心瓶颈在于身份特征嵌入削弱了生成模型的语义表达能力。为此,本文提出一种突破人脸特写限制的IPPG方法,实现身份保真与场景语义创造的协同优化。具体而言,设计了身份-语义分离的双路推理(DLI)管道,解决传统单路径架构中身份与语义表示冲突问题;提出身份自适应融合(IdAF)策略,将身份-语义融合延迟至噪声预测阶段,结合自适应注意力融合与噪声决策掩码,避免身份嵌入对语义的干扰,无需人工掩码;引入身份聚合前置(IdAP)模块,聚合身份信息替代随机初始化,进一步增强身份保真度。实验表明,该方法在非人脸特写场景下仍能稳定高效生成,无需手动标记或微调,可作为即插即用组件快速集成至现有IPPG框架,缓解对人脸特写的依赖,支持电影级角色-场景构建,提升相关领域的个性化生成能力。

原文摘要 · Abstract (English)

Identity-Preserving Personalized Generation (IPPG) has advanced film production and artistic creation, yet existing approaches overemphasize facial regions, resulting in outputs dominated by facial close-ups.These methods suffer from weak visual narrativity and poor semantic consistency under complex text prompts, with the core limitation rooted in identity (ID) feature embeddings undermining the semantic expressiveness of generative models. To address these issues, this paper presents an IPPG method that breaks the constraint of facial close-ups, achieving synergistic optimization of identity fidelity and scene semantic creation. Specifically, we design a Dual-Line Inference (DLI) pipeline with identity-semantic separation, resolving the representation conflict between ID and semantics inherent in traditional single-path architectures. Further, we propose an Identity Adaptive Fusion (IdAF) strategy that defers ID-semantic fusion to the noise prediction stage, integrating adaptive attention fusion and noise decision masking to avoid ID embedding interference on semantics without manual masking. Finally, an Identity Aggregation Prepending (IdAP) module is introduced to aggregate ID information and replace random initializations, further enhancing identity preservation. Experimental results validate that our method achieves stable and effective performance in IPPG tasks beyond facial close-ups, enabling efficient generation without manual masking or fine-tuning. As a plug-and-play component, it can be rapidly deployed in existing IPPG frameworks, addressing the over-reliance on facial close-ups, facilitating film-level character-scene creation, and providing richer personalized generation capabilities for related domains.

个性化生成身份保持场景生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。