arXiv:2601.20511cs.CV2026-01被引 1

用自然语言生成高保真人像合集,保持身份与细节一致

Say Cheese! Detail-Preserving Portrait Collection Generation via Natural Language Edits

  • 通过自然语言编辑参考人像,生成多属性一致的系列照片
  • 构建24K人像合集数据集,支持复杂姿态与视角变化
  • 适合需要批量生成高质量人像的应用场景

随着社交媒体普及,用户对直观创建多样高质量人像合集的需求日益增长。本文提出人像合集生成(PCG)新任务,通过自然语言指令编辑参考人像图像,生成语义连贯的人像系列。该任务面临两大挑战:(1) 复杂多属性修改,如姿态、空间布局和相机视角;(2) 高保真细节保留,包括身份、服装与配饰。为此,我们构建了首个大规模PCG数据集CHEESE,包含24,000个合集、573,000张样本,采用基于大视觉-语言模型的流水线与反演验证生成高质量文本标注。进一步提出SCheese框架,结合文本引导生成与分层身份及细节保留机制,通过自适应特征融合维持身份一致性,并引入ConsistencyNet注入细粒度特征以保障细节一致性。全面实验验证了CHEESE对推动PCG的有效性,SCheese达到当前最优性能。

原文摘要 · Abstract (English)

As social media platforms proliferate, users increasingly demand intuitive ways to create diverse, high-quality portrait collections. In this work, we introduce Portrait Collection Generation (PCG), a novel task that generates coherent portrait collections by editing a reference portrait image through natural language instructions. This task poses two unique challenges to existing methods: (1) complex multi-attribute modifications such as pose, spatial layout, and camera viewpoint; and (2) high-fidelity detail preservation including identity, clothing, and accessories. To address these challenges, we propose CHEESE, the first large-scale PCG dataset containing 24K portrait collections and 573K samples with high-quality modification text annotations, constructed through an Large Vison-Language Model-based pipeline with inversion-based verification. We further propose SCheese, a framework that combines text-guided generation with hierarchical identity and detail preservation. SCheese employs adaptive feature fusion mechanism to maintain identity consistency, and ConsistencyNet to inject fine-grained features for detail consistency. Comprehensive experiments validate the effectiveness of CHEESE in advancing PCG, with SCheese achieving state-of-the-art performance.

人像生成自然语言编辑细节保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。